The GDPR applies in every phase of a generative-AI workflow: scraping, training, inference and logging, while the AI Act only adds system and model obligations on top. Record for each phase who is controller or processor and on which legal basis data is processed.
The trigger is concrete. The European Data Protection Board (EDPB) adopted Guidelines 03/2026 on web scraping in the context of generative AI on 7 July 2026. Shortly afterwards, the Dutch Data Protection Authority (AP) published guidance on generative AI and the GDPR, summarised by Praxikon. Both documents point in the same direction: AI processing of personal data is and remains ordinary GDPR processing, with additional points of attention.
In our estimation, the practical consequence is that organisations can no longer treat the GDPR and the AI Act as two separate compliance files. This is editorial analysis, not a source statement.
What do the EDPB Guidelines 03/2026 say about web scraping for AI training?
According to the EDPB, the scraping of online personal data for the training of generative AI falls under the GDPR. The guidelines set out that principles such as purpose limitation, transparency, data minimisation and accuracy apply before an organisation even considers a legal basis. The EDPB rejects a generic exemption for AI training.
If an organisation uses legitimate interest (Article 6 GDPR) as a legal basis, then according to the EDPB this must pass a three-step test: a legitimate purpose, the necessity of the processing and a balancing against the rights of data subjects. The guidelines were open for consultation at the time of publication, which means the precise wording may still be tightened.
Licentium's legal analysis emphasises that the guidelines apply to three roles at once: parties that scrape themselves, parties that outsource scraping and organisations that buy in ready-made scraped datasets or models. That last group has, according to that analysis, a due-diligence duty: you cannot fully shift the GDPR responsibility onto a data supplier or an external AI lab.
Where does the boundary lie between the GDPR and the AI Act in an AI workflow?
The EDPB topic page on artificial intelligence states that existing GDPR principles — lawfulness, transparency, purpose limitation, data minimisation, accuracy and confidentiality — form the basis for the use of personal data in AI models. The AI Act does not replace these, but comes on top of them.
In our estimation, the dividing line can be summarised as follows:
- The GDPR governs the data processing: which personal data is processed, on which legal basis, with which transparency and with which rights for data subjects — in every phase from scraping to logging.
- The AI Act governs the system and the model: risk classification, technical documentation, risk management, model information and transparency obligations at system level.
Anyone who treats these two tracks as separate files misses the overlap. A single AI action can touch both a GDPR processing operation and an AI Act obligation at once. How you set that up structurally is something we also discuss in our overview of governance tailored specifically to generative AI and in the broader arguments for one governance system for privacy, cyber and AI.
What does it mean that in the Netherlands the AP reviews both the GDPR and the AI Act?
According to Loyens & Loeff's analysis of the Dutch implementation of the AI Act, the AP is designated as market surveillance authority for prohibited AI practices, for the majority of the high-risk systems in Annex III and for the transparency obligations of the AI Act. That supervision runs in parallel to its existing GDPR tasks.
The AP guidance, summarised by Praxikon, explicitly makes generative AI an enforcement priority for 2026 and applies the GDPR rules to both the development and the deployment of generative models. The practical meaning: in the Netherlands one and the same authority looks at your data flows and at the behaviour of your AI system. A strong privacy file alongside a weak AI file — or the other way round — will not hold up under that integral review.
Who bears which responsibility per phase of the AI workflow?
For a concrete workflow, for example an internal knowledge assistant or AI-supported customer communication, the responsibility model can be ordered as follows:
- Developer/provider: GDPR controller for scraping and training (legal basis, special categories, minimisation, transparency) plus AI Act obligations around documentation, risk management and model information.
- Purchasing organisation (controller): its own GDPR responsibility for which personal data ends up in prompts, context and logs, plus AI Act tasks as deployer, such as transparency and human oversight where relevant.
- Processor(s): contractual and technical duties for security, logging and facilitating the rights of data subjects.
- Supervisory authorities (AP and EDPB): review of data flows and system behaviour in conjunction.
Record this division of roles in contracts and in a DPIA. The audit rights and evidence obligations in AI contracts are the instrument here for making the EDPB's due-diligence duty actually enforceable.
How do you demonstrate that your AI workflow is GDPR- and AI Act-proof?
Demonstrability requires documentation that is traceable per workflow. We recommend a control agenda along these lines:
- A data map per workflow: which personal data is processed where and on the basis of which legal basis.
- Documentation of scraping and training provenance: which datasets, with which legal qualification under Guidelines 03/2026.
- Logging per AI session: which prompts, which context sources, which output and which human review.
- A clear division of roles between controller and processor in contracts and DPIAs.
- A link between these elements and the supervisory role of the AP.
Verification tools can support this visibility, but do not change the underlying obligations. A verification layer such as that of IamVera.ai is designed to prepare processing on EU infrastructure and, where the Semantic Privacy Shield is deployed, to send only anonymised content to the selected AI models; if the privacy check fails, nothing is forwarded. That can help to make verification steps and provenance visible, but it is not a guarantee of GDPR compliance. The professional final judgement remains with you. You will find more context in the topic hub on AI privacy and the GDPR.
Sources and references
- Guidelines 03/2026 on web scraping in the context of generative AI
- EDPB Adopts Guidelines 03/2026 on Web Scraping for Generative AI Training
- Dutch implementation of the AI Act: decentralised AI supervision
- Guidance on generative AI and the GDPR (Dutch Data Protection Authority)
- Artificial intelligence | European Data Protection Board
Sources: The article relies on the EDPB Guidelines 03/2026 and the EDPB topic page on AI, the AP guidance on generative AI and the GDPR (via Praxikon) and the analyses by Licentium and Loyens & Loeff.