Blog

AI summaries of long case files: why claim-by-claim checking is needed

Studies from 2025-2026 show AI summaries of long and multiple documents can deviate substantially from the source. Here is how to check them responsibly.

· By

A thin printed summary with highlighted lines lies beside a thick stack of source documents full of coloured tabs, while a hand with a highlighter moves between them.
Every claim in an AI summary should be checked against the source text, because unsupported statements do occur.Image: IamVera.ai — original editorial illustration

AI summaries of long or multiple documents can deviate substantially from the source: research measures hallucination rates from a few per cent to well over 40 per cent, depending on domain and setup. Do not treat such summaries as a quick rendering, but test each claim against the source text and have high-impact sections checked by a human.

The immediate occasion is a study on arXiv on hallucinations in multi-document summarization (updated August 2026). In it, the authors report that in summaries of multiple sources a large proportion of the generated insights are not supported by the source documents, with outliers up to around 45 per cent in the news domain and up to roughly 75 per cent in conversations, including medical consultations. Even more striking: in more than 20 per cent of cases, models produce summaries for subtopics that do not appear in the source at all. For anyone having case files summarised, that is a concrete working risk, not a theoretical caveat.

How much do AI summaries of long and multiple documents deviate from the source?

The figures vary strongly by task and domain. The overview from Inferya on hallucination rates per task type places grounded summarisation benchmarks such as Vectara HHEM (roughly 0.7 to 3.3 per cent for top models) alongside clinical case summaries where, without mitigation, rates up to around 64 per cent are reported. The arXiv study adds the multi-document figures to this.

In our assessment, the most important lesson is not a single figure, but the spread itself: a low benchmark score on short, well-structured texts says little about a long, messy, domain-specific case file. Anyone relying on generic accuracy statistics is measuring the wrong situation.

  • Short, grounded summaries: lowest measured error rates.
  • Long single documents: higher, with error clustering at the end.
  • Multiple sources combined: a substantial proportion not covered by the source.
  • Domain-specific (medical, legal, claims): without mitigation, the highest distortion.

Why do errors accumulate precisely at the end of long summaries?

The research Hallucinate at the Last in Long Response Generation shows that in long outputs hallucinated content systematically shifts towards the end, and that this effect increases as the summary grows longer. Classic average metrics mask this, because the beginning is often correct.

The practical consequence is uncomfortable: professionals often read closing paragraphs and conclusions precisely more cursorily, while that is where the chance of unsupported claims is greatest. Segment-specific checking of the final parts is therefore not a luxury. This connects to broader practical methods for detecting AI hallucinations.

How do I test an AI summary claim by claim against the source document?

A usable framework comes from the peer-reviewed article in Scientific Reports on hallucination detection for summaries. It describes a Question-Sort-Evaluate approach: questions are generated from the summary, sorted by importance and then answered on the basis of the source text. If the answer from the summary deviates from the answer from the source, the claim is suspect. Hallucination is thereby explicitly defined as text that is not supported by the source text.

Translated into daily practice, this amounts to a repeatable sequence of steps:

  1. Split the summary into separate, checkable claims.
  2. Find the supporting passage in the source document for each claim.
  3. Mark claims without coverage as potentially hallucinated.
  4. Give extra attention to the closing parts, where errors concentrate.
  5. Record which claim is supported by which source passage.

Anyone wanting to set this up structurally can combine it with a layered stack for hallucination detection instead of a single tool. More background is available in the topic hub on AI verification and controllable workflows.

Why does human review remain necessary in legal, medical and compliance files?

Even an optimised model does not remove human review. The study in PubMed Central on discharge summaries from electronic health records reports that a fine-tuned model contained at least one factual inconsistency in around 6 per cent of the summaries, and that 94 per cent were found to be clinically acceptable after human verification. Errors persisted particularly in atypical cases. The authors emphasise that human review must remain built in.

In our assessment, that conclusion is more broadly applicable than healthcare: in legal files, insurance claims and compliance reports, AI provides a first draft, but the final decision and final reporting require a demonstrable human checkpoint. In the legal domain, this is elaborated further in a defensible workflow for legal AI research.

What verification should I record before relying on an AI summary?

Reliability only becomes credible when you can show what has been checked. That means recording, per long document, which summary was produced by which model, which claims have been tested against the source, which have been labelled as potentially hallucinated, and where and by whom human review took place. This creates a checkable trail for audits, disputes or incident investigations.

A privacy-focused verification layer such as Vera can support this by routing a task through selected independent models and making verification steps, corrections and sources visible for inspection. That does not promise correctness and does not remove hallucinations; it makes checking possible. The architecture is moreover designed to send only anonymised content onward and works fail-closed: if the privacy check fails, the document does not proceed. The professional final judgement remains with the user. The core stays the same: treat an AI summary of sensitive material as a verifiable sequence of claims, not as a faster reading function.

Sources and references

  1. From Single to Multi: How LLMs Hallucinate in Multi-Document SummarizationarXiv · 2026-08-11
  2. Hallucinate at the Last in Long Response GenerationalphaXiv · 2026-01-13
  3. A hallucination detection and mitigation framework for abstractive text summarizationScientific Reports (Nature) · 2025-12-03
  4. Accurate discharge summary generation using fine tuned large language models in electronic health recordsPubMed Central · 2026-01-17
  5. LLM Hallucination Rate by Task Type (2026 Data)Inferya · 2026-08-21

Sources: The article draws on studies on arXiv and alphaXiv, a peer-reviewed article in Scientific Reports, a clinical study in PubMed Central and a task-type overview from Inferya.

← All articles in this topic ← All articles