Blog

Legal AI still hallucinates in 1 in 5 to 3 in 10 searches: how to build a verification layer

Benchmarks from 2025-2026 show legal AI tools still err in 15-35% of searches. Here is how to set up a verification layer over Lexis and Westlaw.

· By

A wooden desk in two zones: a closed folder with tabs on the left, an open statute book with highlighted case law and pen on the right, and a ticked checklist between them.
Legal AI still hallucinates in 15 to 35 percent of searches; a separate verification layer separates finding from validating.Image: IamVera.ai — original editorial illustration

Do not rely on the tool interface alone: benchmarks from 2025-2026 show leading legal AI tools still hallucinate or misstate the law in 15 to 35 percent of realistic searches. So set up a separate verification layer that separates finding from validation, checks citations and status against primary sources and records human approval.

The trigger is a discussion on the trade blog Artificial Lawyer about whether specialised legal AI tools now need a separate verification layer to stay accurate and relevant in high-stakes files. That question is no longer abstract: recent, independent benchmark research puts concrete figures to it. This article translates those findings into a practical choice for general counsel, knowledge managers and librarians.

How often do legal AI tools hallucinate according to the 2025-2026 benchmarks?

In the study Benchmarking Legal RAG, Stanford RegLab measures hallucination and error rates in purpose-built legal tools such as Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI, alongside a GPT-4 baseline. The outcome: these products hallucinate or answer incorrectly in roughly 17 to 33 percent of the searches examined, even on structured statutory tasks.

AI Vortex's guide AI Hallucination Rates in Legal Tools aggregates recent industry and Stanford data and notes that the accuracy of legal AI tools on complex tasks remains stuck at around 60 to 65 percent. Legal-specific tools therefore perform better than general models, but a substantial share of the answers remains wrong.

An important nuance: these are results from controlled tests with fixed question formulations. In daily practice questions vary more widely; that can increase the error margin in some contexts, although this cannot be inferred directly from these benchmarks.

Why is 'better than a lawyer' not yet reliable enough for high-trust work?

LawNext reports on the Vals AI benchmark, in which both legal-specific and general AI systems reached around 80 percent accuracy and outperformed human lawyers on some research tasks. That is a real result, but it also means the system is wrong in roughly one in five searches.

In our editorial assessment, that is precisely the risk: if a tool is demonstrably better than a junior member of staff, the temptation arises to check less. Particularly in high-stakes work — appeals, supervisory advice, cross-border opinions — a residual error of this size is too large to blindly trust an interface. The recognising outdated AI answers and superseded legislation is part of that too: a convincingly worded answer may rest on out-of-date law.

What exactly does a verification layer over Lexis, Westlaw or comparable legal AI tools do?

LegalAIInsights describes in the article AI in the Legal Field a structured verification protocol over existing tools. Translated into a workable layer, that amounts to the following steps:

  • Separate finding from validation. Use the AI to spot issues and find candidate authority, but do not treat that result as the final answer.
  • Check every citation against a primary source. Does the cited ruling or provision actually exist and is the reference correct?
  • Test the proposition. Does the case genuinely stand for what is attributed to it, or is the meaning distorted?
  • Check the current status. Has the ruling not been overturned, superseded or replaced?
  • Record the approval. Register who verified what, when and against which source.

The verification layer is therefore not an extra tool but a governance workflow that makes visible how an AI finding becomes a checked authority. You will find more background to this thinking in our topic hub on AI verification and controllable workflows.

How do you introduce that verification layer without throwing away your existing tools?

The benchmarks do not argue against Lexis, Westlaw or comparable legal AI tools; they argue for discipline around them. A practical route:

  1. Keep using AI for issue spotting and finding candidate authority, where the speed gain is greatest.
  2. Make a verification protocol mandatory for every authority that ends up in a pleading, memo or piece of advice.
  3. Code that protocol into research checklists and file templates, so that checking is standard rather than ad hoc.
  4. Log per workflow which tool was used, which citations and propositions were verified and who signed off.

This approach aligns with broader controlling legal AI workflows per task and with assessing whether how to check whether AI citations really exist and are correct. The final judgement remains with the lawyer; the layer makes that judgement traceable.

What role can a verification console such as Vera play in this?

The editorial heart of this piece lies not in a product, but in the observation that the figures make a separate, governed verification layer rational. Where such a layer benefits from visibility, a privacy-focused verification console can support this. Vera is not a chatbot and not its own language model, but a verification layer that can route a task through selected independent AI models and make verification steps, corrections, disagreements and sources visible for inspection.

That gives more insight into which tools were used and which citations and propositions were checked against which sources, and where human approval anchored the final result. For sensitive files, the workflow is designed to send only anonymised content to the chosen models; if a privacy check fails, nothing is sent onward. Vera does not promise a correct outcome and does not remove hallucinations — it makes the checking visible, so that the professional final judgement can remain with the user.

Sources and references

  1. Benchmarking Legal RAG: The Promise and Limits of AI Legal Research ToolsStanford RegLab / Deliberative Democracy & Human-Centered AI · 2025-01-13
  2. AI in the Legal Field: Hallucination Risks and a Verification ProtocolLegalAIInsights · 2026-07-28
  3. AI Hallucination Rates in Legal Tools: 2026 DataAI Vortex · 2026-04-17
  4. Vals AI's Latest Benchmark Finds Legal and General AI Now Outperform Lawyers in Legal Research AccuracyLawNext · 2025-10-23

Sources: The article relies on the Stanford RegLab benchmark, the Vals AI benchmark via LawNext, and practical analyses from LegalAIInsights and AI Vortex.

← All articles in this topic ← All articles