The European Data Protection Board (EDPB) has drawn a clear line: web scraping for the training of generative AI falls under the General Data Protection Regulation (GDPR) as soon as personal data are involved. In a press release of 8 July 2026, the EDPB explains the new Guidelines 03/2026, which focus specifically on the scraping of personal data from publicly accessible sources for AI training. The core message is sober but far-reaching: AI training data are not a separate, unregulated category.
What the EDPB guidelines say
According to the EDPB, the core principles of the GDPR — lawfulness, purpose limitation, transparency, data minimisation and accuracy — apply in full to datasets used to train generative AI. The EDPB's consultation page for Guidelines 03/2026 makes the scope concrete: the guidelines are aimed at private parties that scrape personal data themselves or have this done through contractors, and at organisations that reuse datasets scraped by third parties for AI training. Certain parties, such as public authorities, are excluded. An important detail: the text is a draft version and is open for public consultation until 30 October 2026.
A recurring point in the analyses is that public visibility does not in itself provide a legal basis. The legal commentary by Licentium (22 July 2026) emphasises that data do not fall outside the protection of the GDPR simply because they are publicly available on the internet. As a workable basis for scraping for AI purposes, Licentium points to the legitimate interest under Article 6(1)(f), subject to a three-step test. Asking for consent on a large scale is, according to that same analysis, usually not realistic in practice.
The entire lifecycle comes into view
The IAPP pointed out in a contribution of 13 July 2026 that the guidelines run through the full lifecycle of web scraping — collection, storage, training, deployment and deletion — and link each stage to specific GDPR provisions, including Articles 5, 6, 9, 10, 14 and 89. The message is that web scraping for generative AI falls under the GDPR as soon as it concerns identified or identifiable individuals.
TechTimes summarised the practical impact on 9 July 2026 for AI labs and developers: the GDPR applies when it concerns personal data of EU residents, regardless of public visibility. According to that report, the balancing of interests in three steps must take place before scraping, and data minimisation must already happen at the collection stage — not only after the fact. Transparency and accuracy obligations also apply to the data pipeline that feeds AI models.
What stands out across the various sources is that there are no new rules, but rather an interpretation of existing obligations specifically tailored to AI training. For organisations, this means that obligations they already know — a demonstrable basis, a documented balancing of interests, filtering at the source, clear privacy information and mechanisms for the exercise of rights — are now explicitly connected to their AI practice.
AI data flows as verifiable processing operations
The common thread through all the sources is that AI data flows must be treated as processing operations that you can substantiate and demonstrate. Merely naming a legal basis is not enough; the EDPB and the accompanying analyses place the emphasis on demonstrability. Anyone invoking a legitimate interest must be able to show the three-part test. Anyone claiming data minimisation must be able to demonstrate that filtering takes place already at the collection stage. Anyone promising transparency must provide clear privacy information and opt-out options, as Licentium describes.
For professionals who work with confidential information — lawyers, notaries, company doctors, journalists, researchers and compliance teams — the emphasis therefore shifts from "is this allowed?" to "can I demonstrate this?". That is precisely the point where a verification layer can offer support. I am Vera is a privacy-focused AI verification console and not a chatbot or a proprietary language model; the aim is to make control steps visible.
Concretely: the Semantic Privacy Shield is set up so that pre-processing and anonymisation take place on EU infrastructure, and the workflow is designed to send only anonymised content to the selected AI models. If a privacy check fails, nothing is forwarded. This aligns with the EDPB principle that data minimisation and filtering should happen at the source. You can read more about this on the Privacy Shield page.
In addition, multi-model verification makes it possible to place AI answers side by side and to make the underlying steps transparent. Vera does not promise correct or true answers and does not eliminate errors; what it does do is make the check visible, so that the professional final judgement remains with the user. Anyone who wants to view and edit documents within the same secure environment can turn to Vera Office.
What organisations can do now
The guidelines are still a draft version and the consultation runs until 30 October 2026. Even so, the direction is clear enough to start working on your own AI data flows now. Think of documenting the basis and the balancing of interests, building in filtering before data reach a model, drawing up clear privacy information and setting up the exercise of rights. Anyone who treats their AI practice as a regular, demonstrable processing operation is better prepared for the final version of the EDPB guidelines.
The message from the EDPB, Licentium, IAPP and TechTimes comes down to the same thing: AI and the GDPR do not overlap partially but fully as soon as personal data are involved. The distinction between an "AI project" and "data processing" does not exist in legal terms. That makes verifiability — being able to show every step — the core of responsible AI use.