L-MARS: Multi-agent legal system with citation search and auditing

Learn how L-MARS, a multi-agent legal system with citation search and audit, improves accuracy in legal responses.

domingo, 19 de julio de 2026 • 5 min read • Q2BSTUDIO Team

Improving Accuracy in Legal Responses with Multi-Agent AI

In the dizzying advance of artificial intelligence applied to the legal sector, one of the most critical challenges has been to ensure that the answers generated by language models are not only accurate in terms of multiple choice, but also supported by verifiable and traceable sources. The L-MARS (Legal Multi-Agent Retrieval and Scrutiny) system emerges as an innovative proposal that addresses precisely this need: a multi-agent ecosystem that combines active search, automated judicial verification and auditing of each statement against its cited source. This approach represents a quantum leap from classic retrieval and generation pipelines, where citation fidelity is often the weakest link.

To understand its relevance, it is useful to place ourselves in the context of modern legal practice. Lawyers, consultants and law firms handle huge volumes of regulations, jurisprudence and doctrine. The promise of an AI assistant that can answer complex questions with precise references is tempting, but also dangerous if the quotes are made up or come from the wrong sources. This is where L-MARS introduces a paradigm shift: instead of relying on a single generation pass, it deploys a team of specialized agents. A search agent scans legal databases and documents; Another agent, called a judge, examines each atomic statement by labeling it with a taxonomy of six classes (from correct citation to unattainable or unsupported citation). This process of multiple rounds of verification allows you to increase the accuracy of citations significantly.

The results of the first audits on stratified bar exams show a substantial improvement: strict appointment F1 goes from 0.13 in a basic RAG (retake augmented generation) system to 0.25 with the multi-agent judge cycle, and the citation failure rate is reduced from 34% to 13%. In addition, a subsequent step called Faith-Search is introduced, which repairs unattainable citations to below 1%. While it doesn't raise the overall F1, it does ensure that every reference is accessible, a prerequisite in environments where auditability is the law.

Underlying this architecture is a reflection on the nature of AI verification. Traditionally, legal response systems were evaluated with multiple choice success metrics, ignoring whether the cited source actually existed and supported the attributed rule. L-MARS focuses on what experts call 'citation faithfulness', a concept that goes beyond superficial accuracy. For companies developing legal software, this means rethinking the value chain: it is not enough to generate coherent text; each statement must be certified.

From a business perspective, the implementation of multi-agent systems such as L-MARS offers concrete opportunities. Organizations that work with sensitive legal documentation, such as law firms, insurers, or compliance departments, can benefit greatly from these types of automated checks. However, building and deploying such systems requires a solid technological foundation. This is where companies like Q2BSTUDIO can bring their expertise in developing bespoke applications that integrate AI agents, verification flows, and connectors with legal databases. It's not just a language model, but an entire ecosystem where every component—from search to audit—must be designed with surgical precision.

The trend toward autonomous AI agents collaborating with each other is redefining the enterprise software landscape. Instead of monolithic applications, we see modular architectures where each agent fulfills a specific role: one retrieves information, another reasons, another verifies, another generates summaries. For highly regulated sectors such as legal, this modularity allows auditing every step of the process and ensuring traceability. L-MARS is a perfect example of how AI agents can operate together under the supervision of an algorithmic judge, somehow replicating the workflow of a human legal team.

But the application of these concepts is not limited to law. Any area where the veracity and provenance of information are critical—health, finance, compliance—can adopt similar architectures. The key is in the ability to build AI for companies that not only respond, but justify their answers with verifiable sources. And to do this, the underlying infrastructure must be robust. AWS and Azure cloud services provide the scalability needed to run multiple agents in parallel, store legal corpora, and deploy models efficiently. Cybersecurity also plays a critical role: legal data is extremely sensitive, and any verification filters must operate in protected environments. A system like L-MARS, if deployed in production, requires end-to-end protection measures, from encryption to access control.

Another complementary aspect is business intelligence. The results of these citation audits can feed into dashboards that show the reliability of the responses by legal area, by type of source or by model used. Tools like Power BI allow you to visualize the evolution of appointment fidelity over time, identify error patterns, and make informed decisions about agent improvement. In fact, integrating business intelligence services with these multi-agent systems is a growing trend: not only is knowledge generated, but its quality is measured.

Returning to the specific case of L-MARS, a complementary study with 50 questions from the LegalSearchQA set confirms the picture: traditional retrieval and generation pipelines stagnate near an F1 citation of 0.75, while a single agent based on web search falls to 0.22 under external audit. This shows that multi-agent architecture is not a luxury, but a necessity to achieve acceptable levels of trust. The ability to 'hunt' unsupported claims and repair them before final delivery is the hallmark of this type of system.

For companies looking to make the leap to this new paradigm, the way forward is to adopt a tailored software approach that contemplates both agent logic and integration with legitimate data sources. It is not a matter of inserting a generic GPT model, but of designing a workflow where each citation is checked against a curated corpus, and where an algorithmic judge reviews each statement. Technology development companies, such as Q2BSTUDIO, offer precisely this customization capacity: from the selection of the base model to the implementation of verification loops and integration with cloud services. In addition, process automation can take care of repetitive tasks, such as updating legal databases or generating audit reports.

In conclusion, L-MARS represents a milestone in the evolution of AI-based legal systems, but its lesson transcends the legal realm. It reminds us that artificial intelligence cannot be a black box; it must be verifiable, audited and, above all, reliable. The combination of active search, specialized agents and automated judges configures a model that any sector with high standards of informative quality should consider. And to make this a reality, collaboration with experienced technology partners – who are proficient in artificial intelligence as well as cybersecurity, cloud and business intelligence – becomes a differentiating factor. The future of AI-assisted knowledge is not in quick answers, but in verified answers.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.