Blog

The GDPR accuracy principle in generative AI: from hallucination risk to testable workflow requirement

How EDPS and draft EDPB guidance turn AI hallucinations about people into testable GDPR accuracy controls per workflow phase.

· By

A long table split into four stations with stamped papers, folders, a magnifying glass on a highlighted paragraph, and a correction stamp with a page returned to the start.
Accuracy is tested per workflow phase, from sourcing to rectification, with a record of what was checked at each step.Image: IamVera.ai — original editorial illustration

According to the EDPS and the EDPB's draft Guidelines 03/2026, the GDPR accuracy principle covers what a generative AI system outputs about identifiable people, including hallucinated statements, and not only the data it was trained on. Article 5(1)(d) requires personal data to be accurate and, where necessary, kept up to date. Together, the EDPS Orientations of 28 October 2025 and the EDPB draft move the accuracy assessment into the pipeline: from source selection and training data to evaluation and person-related output.

Operationally, that makes accuracy a testable workflow requirement. Each workflow phase needs a control, a pass criterion and telemetry. Four phases carry the requirement: sourcing, training and fine-tuning, output at inference, and rectification.

The pass criteria, telemetry and correction layers proposed below are our operational translation of the texts cited, not regulator-prescribed technical standards.

What does the GDPR accuracy principle require of generative AI?

The accuracy principle requires that personal data processed by a generative AI system are accurate and, where necessary, kept up to date, and that every reasonable step is taken to erase or rectify inaccurate data without delay (Article 5(1)(d) GDPR; Article 4(1)(d) EUDPR for EU institutions). In generative AI this covers more than input data. The EDPS expects accuracy to be safeguarded at every stage of development and use, including the output and inferences a model generates. Controllers must be able to demonstrate compliance: Article 5(2) GDPR, and Article 4(2) EUDPR for EU institutions.

Accuracy is purpose-relative: the EDPB draft describes accurate data as correct and current in relation to the specific purpose they are processed for (paragraph 39).

Scope matters when these texts are cited. The Orientations address EU institutions under Regulation (EU) 2018/1725, which contains the accuracy principle in Article 4(1)(d); they are not binding on controllers under the GDPR. The EDPB guidelines are a consultation draft on web scraping for model training. Both show how supervisory bodies read the principle: as a property the system has to maintain, including its propensity to generate false personal data.

Does the accuracy principle apply to AI output and hallucinations?

Yes, where the output is personal data. The EDPB's draft Guidelines 03/2026 state that the accuracy principle is not limited to collection and training: a model that is expected to produce personal data once it is on the market should comply with it (paragraph 40). The EDPB's ChatGPT Taskforce concluded in 2024, as a preliminary view, that telling users output may be wrong is not sufficient. A hallucinated statement about an identifiable person is therefore an accuracy issue, and a disclaimer alone does not make the processing compliant.

The Taskforce report of 23 May 2024 notes that the probabilistic design of the system can produce biased or invented output and that users are likely to treat output as factually accurate (paragraphs 29 to 31). The report sets out preliminary views, not a final decision.

The draft Guidelines 03/2026, adopted on 7 July 2026 and open for consultation until 30 October 2026, add a design expectation: the controller should take the end result of training into account in order to limit the risk of incorrect output. The EDPS reaches the same point from the deployment side: models trained on representative, high-quality data may still generate false information about people, and an organisation should reassess whether to use a system whose accuracy cannot be sustained.

What is the difference between GDPR accuracy and statistical accuracy?

GDPR accuracy is a legal property of personal data: whether a statement about an identifiable person is correct and current for the purpose it is used for. Statistical accuracy is a performance property of a system: the proportion of its outputs that are correct, measured against test data. The UK's ICO states that the accuracy principle applies to AI inputs and outputs alike, but does not require a system to be 100% statistically accurate. A system can therefore score well statistically and still produce a false statement about one person.

Three notions of accuracy are in play, and each is tested differently:

  • GDPR accuracy (Article 5(1)(d)): correctness of personal data in relation to purpose; tested per statement and per person.
  • Statistical accuracy: share of correct outputs on a test set; tested per model or per release. The EDPS treats such metrics as a useful indicator, not as a data protection measure in themselves.
  • AI Act accuracy (Article 15): an appropriate level of accuracy for high-risk AI systems, with levels and metrics declared in the instructions for use; a system-performance requirement.

According to the ICO, many AI outputs are not meant as fact but as probabilistic estimates about a person; to prevent them from being read as fact, records should mark them as inferences and capture the provenance of the data and of the system that generated them. The guidance applies under the UK GDPR and is under review following the Data (Use and Access) Act.

Read together with the EDPB Taskforce view, marking an inference as probabilistic is expected, but a label does not turn a false factual claim about a person into accurate data. Benchmark scores therefore cannot serve as the pass criterion for GDPR accuracy. The practical test unit is the individual claim about a person, not the model's average.

How do you test accuracy per workflow phase?

Accuracy is tested per workflow phase by attaching a control, a pass criterion and a record to each of four phases: sourcing, training and fine-tuning, output at inference, and rectification. The EDPB draft supplies the sourcing controls: reliable sources, timestamps and validation before training. The EDPS supplies the training and output controls: dataset verification, validation and test sets, provider assurances, and monitoring with human oversight. Rectification closes the loop: an inaccurate output about a person must be traceable, correctable and re-tested. The records per phase make compliance demonstrable under Article 5(2) GDPR.

  1. Sourcing. Use reliable, maintained sources rather than secondary or outdated aggregators, timestamp every record at collection, and validate formats and spot-check samples for factual correctness before training (EDPB draft, paragraph 42). Operationally, the same gate belongs in front of RAG indexing, because retrieved context can flow directly into the output. Pass criterion: every source has an owner, a collection date and a validation result. Telemetry: source registry, collection timestamps, validation log.
  2. Training and fine-tuning. Verify the structure and content of datasets, including third-party ones, hold out validation data during training and a separate test set for the final evaluation, and obtain contractual assurances and documentation such as model cards from providers (EDPS, section 10). Pass criterion: documented dataset curation and an evaluation report per release that includes person-related test prompts. Telemetry: dataset documentation, evaluation reports, supplier assurances.
  3. Output (inference). Monitor outputs and inferences that concern people at runtime, with human oversight (EDPS, section 10), and add a human-in-the-loop verification step scaled to the consequence of an error; the EDPS example is a manual check of every AI-generated CV summary. Confidence scoring can route outputs to review, but a model's self-reported confidence is not evidence of accuracy and should not be the pass criterion (see confidence as a workflow signal). Pass criterion: each material claim about a person is supported by a traceable source, recorded as an inference, or withheld pending human verification. Telemetry: what was checked, against which source, by whom and with what result.
  4. Rectification. Provide an intake route for contested outputs, assess the request, and correct at the layer where the error sits: the record and index in a retrieval system, an instruction or output filter at the application layer, or retraining where the model itself is affected (EDPS, section 14). Pass criterion: the inaccurate claim does not recur across a documented regression set covering the original prompt, paraphrases, relevant context variations and the production model version. Where recurrence cannot be sufficiently controlled, the workflow should block or escalate person-related output in that context. Telemetry: request, assessment, measure, regression results and date.

The structure mirrors the per-phase controls for data minimisation in generative AI: minimisation decides which personal data may enter each phase, accuracy decides whether what leaves each phase is correct. The EDPB draft links the two directly: collecting less makes it easier to keep personal data accurate (paragraph 41). Phase 3 is where a layered approach to hallucination detection and evidence-based hallucination detection at runtime can support compliance controls rather than serve only as quality measures.

What happens when a model outputs inaccurate personal data?

The person concerned can invoke the right to rectification under Article 16 GDPR. The controller must assess the request and, where the personal data are inaccurate, rectify them without undue delay. For generative AI the EDPS distinguishes requests about training data, post-training data, prompts and outputs, and notes that requests about outputs are in principle more common than requests about training data, because outputs affect people directly. The technical difficulty is locating the error: a wrong record in a retrieval database can be corrected and re-indexed, whereas personal data cannot easily be removed from a trained model.

The EDPS works through an example (Orientations, section 14). An institution's HR chatbot produces a false statement about a person's employment history that the institution cannot trace in its internal data. Once the individual supplies the prompts and the full transcript, it treats the request as valid and gives the model an instruction that prevents the same false output from recurring. Operationally, reproduction evidence (prompts, context and output) has to be capturable; without it the controller cannot assess the request.

The appropriate remedy depends on where the inaccurate personal data resides; Article 16 does not prescribe a particular technical mechanism for generative AI systems. A correction can be applied at three layers:

  1. Retrieval layer. Where the false statement originates from a document or record in a knowledge base, correct or delete the source record, then re-embed and re-index it so that stale vectors do not survive the update (see personal data in embeddings and vector databases).
  2. Application layer. Where the source cannot be located, block the specific output with an instruction or an output filter, as in the EDPS example.
  3. Model layer. Models built on the data do not all have to be changed; the exception is a model that itself contains the data or allows it to be inferred, in which case retraining may be needed (EDPS). The EDPB draft notes that, at the present state of the art, removing personal data from a model after training is not straightforward (paragraph 60), and it treats machine unlearning as a longer-term possibility (paragraph 74). For erasure rather than correction, see the right to erasure and model memory.

The reach of the duty for general-purpose chatbots is still being tested. On 20 March 2025 noyb filed a complaint with the Norwegian authority about a ChatGPT output that falsely stated a man had murdered two of his children, attempted to murder the third and been sentenced to 21 years in prison; noyb argues that a disclaimer cannot replace accuracy. The case, in which the Irish DPC is now lead supervisory authority, was still listed as pending on noyb's case page at the time of writing.

What is the role of verification tools such as Vera?

Verification tools do not make output accurate and do not discharge the controller's duty under Article 5(1)(d). Their role is to implement and evidence the output-phase control. Vera is a verification layer, not a chatbot and not a language model of its own: it routes a task through several independent AI models and shows the verification steps, corrections, disagreements and sources. That gives the professional a claim-level basis for review and a record of what was checked. The final judgement on whether a statement about a person is correct remains with the user.

Two mechanisms sit at different points of the inference path, and only one of them concerns accuracy:

  • After generation: the verification chain. Separate roles handle drafting, factual auditing, adversarial challenge and live source checking, and the workflow exposes where reviewers diverged and which uncertainty remained unresolved. This supports the output phase and produces telemetry for the record. It does not establish correctness and does not remove hallucinations.
  • Before the model call: the Semantic Privacy Shield. Vera's privacy architecture is designed so that detected sensitive values are replaced with synthetic session values inside the protected Vera environment before external models are invoked, and only the transformed text that passes the local check is sent; if the check fails, nothing is sent, and the original values are restored locally afterwards. This is pseudonymisation, not guaranteed anonymisation: a data minimisation and confidentiality control.

Sources and references

  1. Regulation (EU) 2016/679 (GDPR), Articles 5(1)(d), 5(2) and 16EUR-Lex · 2016-04-27
  2. Generative AI and the EUDPR: Orientations for ensuring data protection compliance (Version 2)EDPS · 2025-10-28
  3. Guidelines 03/2026 on web scraping in the context of generative AI, Version 1.0 for public consultationEDPB · 2026-07-07
  4. Report of the work undertaken by the ChatGPT TaskforceEDPB · 2024-05-23
  5. What do we need to know about accuracy and statistical accuracy?ICO (UK) · 2023-03-15, under review
  6. AI Act, Article 15: Accuracy, robustness and cybersecurityEuropean Commission, AI Act Service Desk · Regulation (EU) 2024/1689
  7. AI hallucinations: ChatGPT created a fake child murderernoyb · 2025-03-20
  8. Case C097: complaint against OpenAInoyb · filed 2025-03-20

← All articles in this topic ← All articles