Blog

Why pseudonymised AI data stays under the GDPR: the distinction EDPB 02/2026 sharpens

EDPB Guidelines 02/2026 clarify when AI data is truly anonymous and when pseudonymised data stays under the GDPR, with consequences for training and inference.

· By

On 7 July 2026 the European Data Protection Board adopted the Guidelines 02/2026 on Anonymisation. That document sets out in detail when a dataset has been processed far enough to count as anonymous and thereby fall outside the GDPR. For this the EDPB introduces a cumulative test with three criteria: No Record Isolation, No Linkage and No Inference. Only those who pass all three may speak of anonymous data.

For professionals deploying AI on sensitive or high-trust information the practical consequence is immediate: pseudonymised AI data is, on this reading, not a route outside the GDPR. Anyone training models or running inference on hashed, tokenised or encoded data generally remains within the GDPR, with the associated obligations that apply to personal-data processing.

The three-criteria test and contextual re-identifiability

According to Guidelines 02/2026, re-identification risks must be assessed contextually and over time. A dataset is not automatically anonymous because a name has been replaced by a code; the EDPB stresses that anonymity must be tested per relevant entity and over time against the cumulative criteria, including new re-identification techniques.

The EDPB topic page on anonymisation/pseudonymisation summarises the distinction concisely: pseudonymisation is a safeguard that reduces linkability but does not fully break the link to an individual, whereas anonymised data is no longer relatable to an identifiable person and therefore falls outside the scope of EU data protection law. Pseudonymised AI data remains legally personal data; anonymisation is a stricter, qualitatively different status.

An overview analysis by MDP Data discusses the same three cumulative criteria and stresses that classic techniques such as hashing and tokenisation only amount to pseudonymisation as long as re-identification remains technically possible. Pseudonymised data therefore remains personal data under the GDPR and CJEU case law. The editorial reading is that much data labelled in practice as 'anonymous' is in reality pseudonymous and thus falls under stricter obligations.

What this means for training, inference and publication

The analysis by SecurePrivacy explicitly links the three-criteria framework to AI contexts and the web scraping guidelines: anonymity is assessed entity-relative and organisations must continue to see pseudonymisation as processing of personal data that falls under GDPR rules. Pseudonymisation during data collection and AI training is a security measure, not a way out of the GDPR. This piece also connects pseudonymisation to Guidelines 01/2025 on pseudonymisation as a separate technique.

A sector case makes this concrete. The analysis by Iliomad Health Data describes how Guidelines 02/2026 must be applied in clinical trials: anonymisation requires a cumulative three-criteria test and contextual risk analysis per relevant entity, while pseudonymisation, such as re-coded subject IDs, remains subject to Article 4(5) GDPR and requires full GDPR compliance, including key management and limitation of re-identification paths. AI-driven analyses and model training on pseudonymised trial data effectively remain processing of personal data and therefore call for DPIAs, contracts and technical safeguards.

For high-trust professionals this means that pseudonymisation and anonymisation no longer count as interchangeable privacy labels. They are two different design and governance categories. Anyone setting up AI workflows must explicitly model which data layers are deliberately kept pseudonymous as a security safeguard within personal data, which datasets are truly anonymous after a rigorous three-criteria test and may therefore be shared or published differently, and where re-identifiability paths and key management must be recorded.

A verification layer such as IamVera.ai can help make that dividing line visible per workflow: which datasets are pseudonymous and which are anonymous, which re-identification scenarios and tests have been carried out, and where governance and audits should focus. The processing is designed so that pre-processing and privacy protection take place on EU infrastructure and so that the workflow strives to send only content that, according to the internal check, is treated as anonymised to the selected AI models; when a privacy check fails, nothing is sent onward. That gives more insight into the question of whether any 'pseudonymous' AI processing has been treated as anonymous by mistake. The professional final judgement remains with the user.

← All articles