Automated hallucination detection shifts in 2026 from general model assessment towards evidence-based runtime verification: systems break answers down into individual claims and check these against external sources or formal policy rules. AWS, Collibra and TrustScale show this, but their accuracy claims are no substitute for independent, domain-specific evaluation.
On 23 June 2026 AWS announced automated refinement workflows for Automated Reasoning checks in Amazon Bedrock Guardrails. These checks use formal logic to test generative answers against a customer-defined policy, flag policy-related hallucination risk and provide verifiable explanations; they do not by themselves establish that an answer is true against external reality. For anyone working with sensitive information, this can offer an additional, checkable step: answers are then not only judged on plausibility, but also tested against explicit policy, with explanations about the outcome. That does not replace source checking or human review.
In our assessment this is the most important change: the question shifts from does this sound plausible to against which rule or source has this been tested, and with what evidence. That makes detection more usable, but shifts the burden to the quality of the policy and the sources themselves.
What exactly does hallucination detection measure?
Hallucination detection is not a single measurement point. It can test several different things:
- Factuality: whether a statement is correct against external reality or a reliable source.
- Faithfulness: whether the answer stays true to the supplied context or documents, without additions.
- Policy compliance: whether the answer meets explicitly defined, formal rules.
These three do not coincide. An answer can be faithful to a faulty source document, or factually correct but in conflict with policy. Anyone deploying a detector must therefore know which of these three the system actually checks. The AWS announcement focuses on policy compliance through formal logic; that is something other than checking against broader reality.
Which automated detection methods exist?
For a practical overview, the methods discussed here can be organised into four working families. They differ fundamentally in what they make plausible and where they fail:
- External claim and source verification. The answer is broken down into atomic claims, external sources are gathered and each claim is validated separately. TrustScale describes Argus as such a black-box system: ingestion, research, validation, comparison, scoring and protection, with colour-coded signals, source references and inline correction. The weak spot lies in source selection and retrieval: a wrongly chosen or outdated source leads to a wrong judgement.
- Formal logic and policy checks. Answers are mathematically tested against explicitly recorded rules, as in the Automated Reasoning checks from AWS. The quality of this validation depends to an important degree on the quality, completeness and unambiguity of the policy; AWS therefore describes iterative policy improvement and the removal of ambiguity.
- Context and knowledge-graph methods. Here hallucination risk is linked to curated context and graphs. In its announcement of new capabilities, Collibra describes a context graph (Live Map), governance agents (Maestro) and Guardian Agents with machine-readable Agent Contracts for runtime oversight.
- Model- or agent-based assessors. Another model assesses the output. The risk here is that the assessor can reproduce the same errors as the model being checked.
Why is no single detector universally reliable?
Each family has a structural dependency that can undermine the detection. That is not a detail, but the core of the limitation:
- Formal checks are only as good as the policy that underlies them.
- Source-based systems stand or fall with source selection and retrieval.
- Model-based assessors can share the blind spots of the underlying model.
On top of that, the published accuracy figures are vendor claims. TrustScale cites figures about accuracy and speed; these do not constitute independent validation. The survey figures in the Collibra post have likewise been presented by the vendor. In our assessment the most important open question is therefore how these systems perform in independent, domain-specific tests. As long as these are missing, a high score on a vendor test remains a weak foundation for business-critical deployment. That aligns with the broader line of layered verification instead of standalone tools.
How do you set up detection for sensitive work?
The practical message for professionals with confidential information is a layered workflow, not a single product. We advise the following steps, as editorial analysis and not as a source statement:
- Break the answer down into individual claims instead of assessing the whole at once.
- Check each claim against the relevant external source or against explicit, formal rules.
- Preserve the underlying evidence, so that a judgement can be reconstructed later.
- Route uncertain or conflicting results to human review.
- Record the model version, the policy used, the sources consulted and the correction decision.
This workflow is a practical editorial recommendation for validation; see also the four measures NIST links to validation of generative AI. Anyone deploying agents should also involve bounding agent actions at runtime, precisely the area at which Collibra's Agent Contracts are aimed. An overview of methods and products is in the hub AI verification: methods and products.
The product developments at AWS, Collibra and TrustScale make hallucination detection concretely runtime- and evidence-based. That is a gain. But the absence of independent, domain-specific evaluation remains a structural risk: organisations must set up their own validation and monitoring processes alongside the tools on offer, not instead of them.
Sources and references
- Automated Reasoning checks in Amazon Bedrock Guardrails add new policy refinement workflows
- Collibra Launches New Capabilities to Reduce the Hallucination Tax on Enterprise AI
- TrustScale Launches Argus to Detect and Correct AI Hallucinations and Power Safe Enterprise AI Adoption
- Argus — AI Hallucination Detection & Output Verification
Sources: The article draws on the official AWS announcement about Automated Reasoning checks in Amazon Bedrock Guardrails and on Collibra's and TrustScale's own posts.