Never trust an AI citation on formatting alone: check every source in three steps, namely whether the source exists, whether the URL or DOI truly points to that source, and whether the source supports the specific claim. Let AI only propose candidate sources and test existence and content independently, because models judge their own citations weakly.
A study on arXiv, published on 3 April 2026 under the title Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, examined 53,090 URLs from deep-research and search-augmented systems and 168,021 URLs across 32 fields. The researchers found 3% to 13% presumably fabricated URLs and 5% to 18% non-resolving URLs. Deep-research agents produced more citations, but hallucinated URLs more often than search-augmented models. More source references therefore does not automatically mean more reliability.
That is the core for the reader: a long, neatly formatted reference list is not proof. Anyone working with sensitive or high-trust information should set up source checking as a separate, verifiable workflow rather than relying on the model that wrote the text. This piece sets out the recent findings and translates them into a concrete verification chain. Where we draw an inference that does not appear literally in a source, we mark it as an editorial assessment.
Why do even deep-research agents not deliver reliable source references?
Deep-research agents are designed to search, read and summarise independently, and in doing so produce more citations than simple search-augmented models. According to the same arXiv study, this comes with a higher share of hallucinated URLs. The researchers do report that an open-source checking tool reduced the share of non-resolving URLs in their experiments to less than 1%, but that outcome depended on correct tool use.
In our assessment, the most important lesson from this is that the volume of references says nothing about their quality. An agent that supplies dozens of sources may at the same time introduce more errors than a model that names only a few. The correction lies not in the generation, but in a separate checking step afterwards.
Which three types of citation error should I check separately?
The sources show that 'a faulty citation' is not one problem but several, each requiring its own check. Based on the studies, we distinguish three layers:
- The source does not exist. The model invents authors, titles, journals or years. The study GhostCite (arXiv, 14 May 2026) tested 13 language models with 375,440 generated citations and found hallucination rates of 14.23% to 94.93%.
- The metadata is wrongly linked. A valid DOI or URL points to an existing but entirely different article. The editorial statement On generative AI, fabricated references and ethical publishing in the Journal of Science Communication (1 May 2026) describes precisely such 'ghost references' with DOIs that lead to a different work.
- The source does not support the claim. The source exists and the link is correct, but the text does not say what the AI asserts. This is the subtlest error, because form and provenance then appear entirely in order.
A check that looks only at form catches at most the first layer. The second and third layers require someone to actually open the source and read the relevant passage.
Why can an AI model not reliably judge its own citations?
An obvious answer would be: let the model check its own sources. The GhostCite study measures exactly that assumption. According to the researchers, language models were on average only 38% accurate when they had to judge for themselves whether references were valid. The GhostCite study also reports that 1.07% of papers contained at least one invalid citation.
That makes self-verification unsuitable as the sole safeguard. If the checker is just as fallible as the generator, apparent certainty arises instead of control. This aligns with broader findings on the limits of self-verification by a single AI model. External checking of both the existence of the source and the support for the claim is therefore, in our assessment, not extra diligence but a minimum requirement.
What does a verifiable verification chain for sources look like?
On the basis of the studies, a verification chain can be built that addresses the three error layers separately and records the human decision. We propose the following order as editorial advice:
- Let AI only propose candidate sources. Treat every reference as a proposal, not as proof.
- Verify existence. Check author, title and year via an independent register, catalogue or archive, separate from the model that gave the citation.
- Test the linking. Open the DOI or URL and confirm that it leads to exactly the stated publication. In doing so, distinguish outdated links (link rot) from fabrication; only the former can be restored.
- Read the relevant passage. Tie each claim to the exact sentence or table that supports it, not to the source as a whole.
- Record exceptions and approval. Set down which sources remained unclear and who took the final decision.
The distinction between link rot and fabrication matters because a non-resolving URL is not automatically invented. Anyone who treats citations as separate objects rather than as a text fragment can make these steps traceable per source. More on this type of checking can be found in the topic hub on AI verification and source checking.
What does this mean for work with sensitive or high-trust information?
In consequential decisions a faulty source undermines not only the text, but also the knowledge base beneath it. A review described by CIDRAP (University of Minnesota, 16 April 2026) found 4,046 presumably fabricated references among 97.1 million checked references from biomedical publications. According to the researchers, the share of papers with at least one such reference rose from roughly one in 2,828 in 2023 to one in 458 in 2025 and one in 277 in the first seven weeks of 2026. The described references were often specific in content, correctly formatted and attributed to real researchers.
It is precisely that plausibility that makes superficial checking inadequate. A citation that looks professional is no proof that the cited publication substantiates the claim. For anyone delivering high-trust work, this means that source checking should be an explicit step in the verification layer for high-trust decision-making, with recorded exceptions and a human final judgement. The studies provide the error rates; responsibility for the decision remains with the professional.
Sources and references
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
- GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
- On generative AI, fabricated references and ethical publishing
- Review uncovers rising rate of fake references in published biomedical papers
Sources: The article draws on arXiv studies (GhostCite and the deep-research reference study), an editorial statement in the Journal of Science Communication and a review described by CIDRAP of the University of Minnesota.