Blog

AI hallucinations: from vague error to measurable risk in decision-making

Recent benchmarks and documented court cases show that AI hallucinations are highly context-dependent. What does this mean for professionals?

July 21, 2026 · Victor Angelier

AI hallucinations are no longer an abstract caveat about the use of language models. Recent academic research and documented incidents show that they constitute a measurable, highly context-dependent risk that feeds directly into professional decision-making. The figures reveal something uncomfortable: a model that almost never gets it wrong in a tightly defined test can, in practice, regularly produce incorrect or poorly substantiated answers.

What the benchmarks show

The 2026 AI Index Report from the Stanford Institute for Human-Centered AI devotes explicit attention to hallucinations in its chapter on Responsible AI. Despite progress, large language models continue to produce factual inaccuracies, while demand for these systems in high-risk domains such as law and medicine is increasing. According to the report, hallucination rates across 26 frontier models vary widely — from roughly a quarter to almost all answers — depending on how a task is formulated. Hallucinations have thus become a situational, yet persistent risk.

The year before, the AI Index Report 2025 from Stanford HAI documented that frontier models can achieve very low rates on certain grounded accuracy benchmarks. For models such as GLM-4-9b-Chat, Gemini-2.0-Flash-Exp, o1-mini and GPT-4o, values around 1.3 to 1.5 per cent were measured on some tests. At the same time, that same chapter describes well-known incidents in which lawyers submitted court filings with fabricated citations that had been generated by language models. That contrast is at the heart of the problem: a good score on a limited test says little about behaviour in a real workflow.

Legal practice as a test case

Nowhere is that difference so sharply documented as in the law. The study AI on Trial from Stanford RegLab and HAI concluded that general chatbots hallucinate between 58 and 82 per cent of the time on legal questions. Specialised legal tools from LexisNexis and Thomson Reuters performed better, but still hallucinated somewhere between 17 and 34 per cent of the time. Importantly: systems that use retrieval augmentation — which link answers to retrieved sources — do reduce hallucinations, but do not eliminate them. And models often proved unaware of their own errors in legal reasoning and citation.

A review article from Kallam.ai brings together recent empirical findings and field reports. It states that in 2025 there were more than 200 documented cases of AI-fabricated citations reaching a judge, and at least 66 judicial sanctions for the misuse of generative AI, with fines running to tens of thousands of dollars. These figures from a non-peer-reviewed source stand in stark contrast to the optimistic benchmark scores, and illustrate that hallucinations have tangible consequences for decisions and for the trust that professionals enjoy.

A structural problem, not an incident

Why can the problem not simply be trained away? The peer-reviewed review study in EMNLP Findings, A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models, defines hallucinations across different types of foundation models and describes them as a major obstacle to deployment in domains that require high reliability. The authors emphasise that hallucinations are structurally linked to the way these models are trained and used — they are not rare aberrations, but a property of the system. Detection and mitigation methods exist, but none of them fully resolves the problem.

For those who make decisions based on confidential or highly sensitive information, the conclusion is sobering: the apparent self-assurance of a single model is not a reliable indicator of accuracy. An answer that sounds fluent and convincing may equally contain a fabricated source or flawed reasoning.

What this means for professionals

The reasonable response to this picture is not to avoid AI, but to treat AI output as an unconfirmed claim that requires checking. That calls for working methods in which verification is visible: comparing answers across multiple models, linking statements to verifiable sources, and a checkable trail of what was asked and answered. Human oversight remains the final safeguard here — the professional final judgement rests with the user, not with the model.

At this point the role of a verification console such as I am Vera comes into view. Vera is not a chatbot and not its own language model, but a verification layer that makes the checking steps around AI answers visible. Through multi-model verification, Vera can help reveal where models diverge from one another — precisely the kind of signal that is useful with hallucinations. That does not make hallucinations impossible, but it provides more insight into the reliability of an answer than blind trust in a single system.

Because professionals often work with sensitive documents, the architecture is designed accordingly. With the Semantic Privacy Shield, pre-processing and anonymisation take place on EU infrastructure; the workflow is designed to send only anonymised content to the selected AI models, and if a privacy check fails, nothing is forwarded. Documents can be viewed and edited within the same secure environment via Vera Office, and the verification steps can be recorded as a checkable evidence trail.

The message of the recent research is ultimately simple: hallucinations are not a residual error that disappears of its own accord as models improve. They are a systemic reliability risk that calls for structural safeguards. For work in which a wrong decision is costly, that means: checking, cross-checking and documenting — instead of trusting the tone of the answer.

← All articles