The European draft standard prEN 18229-1 structures AI event logging, but sets no universal retention periods and prescribes no concrete hallucination detection. Organisations must set up claim checking at claim level themselves, log evidence and decisions, and choose a risk-based retention per log category based on the EU AI Act, GDPR and incident risk.
On 17 July 2026 the Slovenian Institute for Standardization published the draft standard oSIST prEN 18229-1:2026, AI trustworthiness framework – Part 1: Logging. It is a draft standard that describes terminology, requirements and guidelines for the event logging of AI systems. According to the European standards repository STANDICT, the standard was still in an approval or consultation phase in September 2026 and is not a final, binding standard. The practical consequence: the standard helps determine what you log, but the checking of the correctness of AI answers and the retention periods have to be filled in yourself using other frameworks, such as the EU AI Act (Regulation (EU) 2024/1689) and the General Data Protection Regulation (GDPR).
What must an AI event log record in practice according to prEN 18229-1?
The draft standard treats logging as a distinct component within the AI trustworthiness framework. A usable event log proves more than that an answer appeared. In our estimate, these are the fields that a governance log for a sensitive workflow must connect:
- which model and which model version were active when the answer was generated;
- which input, prompt and sources were used;
- which individual claims from the answer were checked;
- which warnings, anomalies or uncertainties arose;
- who carried out the verification and which governance decision followed from it;
- the timestamp, the reason for the change and the approved version.
Decoupling these fields makes a reconstruction possible. Without claim and decision context a log remains a collection of individual answers without a demonstrable chain. In our topic hub on AI governance and controllability this distinction between recording and accountability recurs more often.
Which retention periods apply to AI logs under the EU AI Act and GDPR?
The draft standard itself sets no universal retention period, and it would be incorrect to derive one from prEN 18229-1. The standard structures the event logging; the retention period per log category follows from other frameworks. In our estimate, the following factors determine how long you keep a specific log category:
- applicable legislation and sector rules, ; any EU AI Act obligations must be assessed for the specific system and log category
- the purpose for which the log was created and its purpose limitation;
- the incident risk and the chance that an event later needs to be reviewed;
- privacy requirements from the GDPR, including data minimisation for logs containing personal data.
The practical choice is therefore not a single figure, but a classification per category. Performance measurements, model decisions, warnings and verification outcomes may be assigned different periods. Anyone who wants to be able to demonstrate human oversight of a high-risk decision keeps the decision context available longer; we wrote earlier about demonstrating human oversight of AI decisions via logging.
How does hallucination detection at claim level work in AI systems?
The academic literature shows that detection becomes more reliable when you split an answer into atomic, verifiable claims and test them separately. In the research FactScore, published via the Association for Computational Linguistics, that checking proceeds in four steps: claim decomposition, finding evidence, assessing evidence and determining the precise hallucination location. This makes it possible to know not only that an answer is wrong, but also where.
The study MARCH: Multi-Agent Reinforced Check for Hallucination, also via the Association for Computational Linguistics, stresses that the checker must test the claims independently without taking the original answer as its starting point. The study identifies confirmation bias in systems where a model assesses its own output. In our estimate, this is an important operational complement to prEN 18229-1: the draft standard structures event logging, while organisations can use methods such as MARCH and FactScore to document how generation and checking are separated. We work out a layered approach in our analysis on the layered verification workflow against hallucinations.
A second warning comes from the research Detecting Hallucinations in Authentic LLM–Human Interactions. That study introduces AuthenHallu, based on real human-model dialogues, and reports that standard LLMs as detectors still perform insufficiently in realistic situations. Performance on artificial benchmarks therefore does not automatically apply to real FAQ interactions. Domain risk and human escalation remain part of the process.
Which governance steps follow after verifying an AI FAQ?
A verified FAQ is, in our estimate, not an end point but a versionable governance object. The following steps connect to the log fields from prEN 18229-1 and to the claim checking from the research:
- mark the FAQ as confirmed, uncertain, corrected or withdrawn;
- keep the source passages used and the detection scores per claim;
- route uncertain or contradictory claims to a competent human reviewer;
- formally approve the new version before publication follows;
- record the reason for the change, the responsible person, the model version and the date;
- publish only with a version number and plan periodic re-verification when sources, policy or models change.
Checking the underlying references is part of this; we described that process in checking AI citations in three steps. The common thread: the combination of the European logging development and the recent detection studies shifts the starting point from AI has answered to a demonstrable chain of event, evidence, judgement and governance action.
Sources and references
- oSIST prEN 18229-1:2026 - AI trustworthiness framework - Part 1: Logging
- AI trustworthiness framework - Part 1: Logging (CEN/CENELEC prEN 18229-1)
- MARCH: Multi-Agent Reinforced Check for Hallucination
- Detecting Hallucinations in Authentic LLM–Human Interactions
- FactScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Sources: The article draws on the draft standard prEN 18229-1 from SIST and STANDICT and on research from the Association for Computational Linguistics (MARCH, AuthenHallu and FactScore).