With the adoption of Guidelines 03/2026 on web scraping in the context of generative AI on 7 and 8 July 2026, the European Data Protection Board (EDPB) set out in detail how the GDPR may apply to large-scale scraping and the training of AI models. The guidelines have been released for public consultation and, together with Opinion 28/2024 and the new anonymisation guidelines (Guidelines 02/2026), form a clear signal: an abstract appeal to 'AI innovation' is no substitute for the ordinary obligations under the GDPR.
For professionals who work with confidential or high-trust information, that is relevant. As soon as an AI model works with potentially identifiable personal data, lawfulness, purpose limitation, transparency and data minimisation simply continue to apply. This article sets out what the new lines mean and where the overlap, the differences and the responsibilities lie.
Three myths the EDPB debunks
The Guidelines 03/2026 and the earlier Opinion 28/2024 touch on three persistent assumptions.
Myth 1: AI models by their nature fall outside the GDPR. In Opinion 28/2024 the EDPB establishes that AI models are, as a rule, rarely truly anonymous. Through model queries, meaningful inferences or even re-identification can occur. As long as that is possible, the model remains within the scope of the GDPR.
Myth 2: publicly available data may be scraped without limit. The EDPB states explicitly in Guidelines 03/2026 that 'publicly available' is no free pass and that web scraping involving personal data remains subject to the GDPR. Web scraping involving personal data always falls under the GDPR. As the analysis by Alston & Bird in EU Regulators Outline GDPR Requirements for AI Web Scraping describes, organisations that scrape themselves or purchase pre-scraped datasets must document a legal basis and a legitimate interest, apply data minimisation and take account of the duty to provide information.
Myth 3: synthetic or 'anonymised' models automatically fall outside data protection law. The anonymisation guidelines introduce a three-step test: No Record Isolation, No Linkage and No Inference. The analysis by Secure Privacy in the new 2026 anonymisation test shows that models and synthetic data often remain within the GDPR, precisely because memorisation and inference attacks make re-identification possible.
Where AI and the GDPR reinforce each other
The core principles of the GDPR are directly applicable to AI training and use. That means recording, per AI workflow, who is the controller, which personal data are used when (training, fine-tuning, inference or logging) and on what legal basis this happens. Opinion 28/2024 clarifies that the use of 'legitimate interest' as a basis for training must be substantiated considerably more strictly, with a demonstrable balancing of interests.
The EDPB moreover points out that models trained on unlawfully processed personal data can affect downstream processing and that organisations must carefully assess the provenance of training data. Anyone using third-party models or datasets therefore cannot escape due diligence on the provenance of the training data.
Where AI imposes additional burdens
The difference from classic databases lies in the model memory itself. With traditional storage, personal data are traceable in records; with AI, part of the risk shifts to what the model has 'memorised'. The three-step test forces organisations to assess not only datasets, but also model architecture and the release of models before they label a model as anonymous.
The practical consequences become concrete in the compliance guide GDPR for AI Systems by Strac, which works out AI and GDPR compliance in practical terms in line with the EDPB guidelines and the EU AI Act. From this follows a governance picture: per AI system an inventory with a lawfulness basis, a combined DPIA and FRIA, registration in the record of processing activities (ROPA) and audit logging of AI interactions in high-trust workflows.
From legal text to operational governance
Translating a guideline into practice requires visibility: over which models and datasets run where, which GDPR roles (controller, processor or joint controllers) belong to which processing, which legal basis and balancing of interests have been recorded, and how prompts, outputs and logs are available for audits.
At that point a verification layer can offer support. IamVera.ai is not a chatbot and not its own language model, but a privacy-focused verification layer for professionals who work with confidential information. Vera can route a task through selected independent AI models and make verification steps, corrections, disagreements and sources visible for inspection. This supports review and control and gives more insight into what happens; it is no guarantee of correctness and it does not automatically reduce hallucinations.
For the anonymisation question, the Semantic Privacy Shield is relevant. It can replace sensitive document values before AI processing with synthetic, session-only equivalents on EU infrastructure. The AI chain then analyses the synthetic version, after which the original values can be restored locally. The architecture is designed to send only anonymised content onward, and the workflow is fail-closed: if the privacy verification fails, the document is not sent onward. That is an architecture choice, not a promise about the degree of anonymisation or about full GDPR compliance.
In the verification and log function those steps become inspectable, which can help during audits to see which verification steps and sources have been used. The professional final judgement always remains with the user.
What this means for your AI landscape
The new EDPB lines compel three concrete actions. First: record, per AI workflow, the role and legal basis for training, inference and logging. Second: demonstrably test whether models and datasets fall outside the GDPR via the three-step test, and otherwise arrange full GDPR compliance. Third: make responsibilities visible across models, suppliers and data flows. AI compliance under the GDPR is therefore, in 2026, no longer a paper exercise, but a verifiable layer that must demonstrably hold up.