A compact summary of a thick dossier looks reassuring: the model has reduced a hundred pages to a manageable text, with tidy scores on the usual quality metrics. But several recent 2026 ACL/Findings and arXiv studies suggest that this reassurance can be misleading in long-document settings. Taken together, the studies suggest that the evaluated long-document models can miss key claims, mix up source details, or introduce unsupported facts — while several common factuality metrics were less reliable in long-document settings and could give inconsistent scores, although robustness varied by metric.
The occasion for this piece is PROBE: PROcess-Based BEnchmark for Hallucination Detection. This work treats summary reliability as more than a single final score. Instead, summarising is explicitly treated as a multi-step verification problem: breaking down claims (claim decomposition), finding the right source evidence, assessing that evidence, and localising hallucinations. The study reports that current models struggle in particular with finding correct source passages and localising hallucinations. We interpret this as making evidence finding and hallucination localisation especially important verification steps.
Classic metrics become less reliable with long documents
That process view is no luxury. In Stress Testing Factual Consistency Metrics for Long-Document Summarization researchers test six common, reference-free factuality metrics on long-form benchmarks in fiction, legal and scientific domains. The paper describes that these metrics give inconsistent scores for summaries that are substantively equivalent, and that they prove less reliable in particular for information-dense claims that closely resemble several source passages, although robustness varied by metric.
For practice, this points to something uncomfortable. A report, ruling or policy memorandum of dozens of pages can receive a neat, compact summary with high scores, while crucial exceptions or shifts in context have unnoticeably disappeared. Several of the metrics that are often used to determine whether a summary is "good enough" prove less stable in long-document contexts and can assess summaries with equal content differently. For practice, this suggests, in our estimation, that anyone who relies on such scores may underestimate factual errors.
From a single summary to verifiable claims
Set against this problem is a growing range of techniques for detecting hallucinations at the claim level. Hallucination Detection in Long-Form Text Generated by LLMs introduces the LHD benchmark and an approach with a hyper-relational knowledge graph (HRKG-HD). It proposes an HRKG-HD approach that uses multi-hop relational reasoning to detect hallucinations in long, fact-rich texts. The benchmark makes key parts of long-form hallucination detection concretely measurable: long-form outputs can be broken down into separate facts and relations, which are then held against source segments. In our analysis, this means checks can focus on individual claims rather than trusting or distrusting an entire text as one block.
Multimodality does not solve the problem by itself either. MMLDSum-LLM presents, with MMLDSum-Bench, a benchmark in which text and image occur together, across different domains and context lengths. In the evaluated settings, the benchmark found that multimodal models still showed these coverage and cross-modal factual-consistency problems, even with large context settings. In our analysis, this points to the desirability of additional verification layers for aligning image and text in dossiers that combine text, tables and illustrations.
Normative documents are hit hardest
The study Summarising Regulations: An Empirical Study of Long-Document Summarisation Methods under Extreme Compression discusses this in a concrete domain. The paper reports that, for regulations under extreme compression, important information can be lost or shifted. Retrieval-augmented variants do not always perform better than classic internal-structuring and semantic chunking methods. The research reports that under extreme compression, important normative details — including exceptions, conditions and definitions — may be lost or altered.
That underlines why a summary in itself is not a reliable substitute for the source text. For practice, this suggests, in our estimation, that article- and clause-level verification is sensible for contracts, legislation, and compliance reports.
Summaries as provisional narratives
The common thread through these sources is, in our analysis, the same: with long documents reliability is not a single score, but a process that should be visible and traceable. On the basis of these studies it is, in our estimation, sensible to treat AI summaries as provisional narratives rather than as completed dossier documents. In our view, a robust workflow would make clear which source parts have been summarised, how each claim can be traced back to a paragraph or article, which hallucination detection has been applied and where human review remains necessary.
Here lies, in our estimation, the role of a verification console such as I am Vera. Vera is not a language model and not a chatbot, but a verification layer for professionals working with confidential or high-trust information. The workflow is designed to pre-process and anonymise documents on EU infrastructure before content is offered to selected AI models via the Semantic Privacy Shield; the workflow is fail-closed, so if that privacy check fails, nothing is forwarded. In our analysis, by letting a task run through several selected models and making verification steps visible, Vera can help users treat summaries as a verifiable intermediate product rather than as an endpoint.
Vera does not remove the risk of hallucinations and does not make any assurance about correctness or truth. What it does support is more insight into the chain: which source, which claim, which check. In our analysis, the 2026 research supports the case for more visible verification in long, sensitive documents. Professional judgement and the final decision remain with the user, who may accept the summary after the claims have been checked against the source.