Blog

Right to erasure now reaches into model memory

EDPB guidelines on web scraping and unlearning research show that deletion requests in 2026 reach into model memory and demand demonstrable policy.

· Victor Angelier

During its plenary meeting of 8 July 2026, the European Data Protection Board adopted the Guidelines 03/2026 on web scraping in the context of generative AI and opened a public consultation on them. The guidelines make clear that web scraping for generative AI falls under the GDPR as soon as personal data is processed in the process. A technically relevant passage that the EDPB includes in the guidelines reads: once the model is trained, personal data cannot be easily deleted from a model. With this, the discussion about deletion requests shifts, in our assessment, from the training database to the model itself.

For organisations that deploy generative AI, this means, in our assessment, that retention and deletion policy is no longer merely database management. It is, in our analysis, an architectural question surrounding model memory. Below we set out what the sources describe and which workflow logically follows from it, in our analysis.

The right to erasure reaches into the model

The SPE report commissioned by the EDPB, Effective implementation of data subjects' rights, examines legally and technically how rectification and erasure can be applied to AI systems trained on personal data. The report discusses that effective erasure in AI can also affect the influence of training data on the model, and explores retraining and unlearning as approaches for this. The report describes full retraining with excluded data as the most effective and complete known way to reduce the influence of specific data from a model. It characterises machine unlearning techniques as relatively young, approximate approaches in development, with cited analyses pointing to possible privacy and bias risks.

The legal analysis GDPR and AI: The "Right to Be Forgotten" Now Means Unlearning translates this into supervisory practice. The analysis states that models should not automatically be regarded as anonymous, and discusses that supervisory authorities, in serious cases, also see model deletion as a possible measure. For individual requests, the analysis outlines a practical workflow with impact analysis on models and possibly planned retraining.

The media source EDPB blocks AI firms from using consent as an excuse to scrape works out the consequences of Guidelines 03/2026 for the sector. The media source discusses the same technical difficulty: once trained, models are not easily cleansed of personal data. In our analysis, that acknowledgement actually raises the bar, because controllers must then be able to show which measures they take beforehand and afterwards.

The technology lags behind expectations

At the same time, 'forgetting' is not yet a technically solved problem. The report Google Research validates an audit test, but not yet on LLMs reports that recent unlearning procedures on large models leave residual imprints, and that audit tests for 'forgetting' have largely been validated on synthetic data and smaller models, not on large language models. Together, these sources suggest that current unlearning techniques and audit tests, as described, do not yet provide robust certainty that deleted data no longer influences model outputs in the context studied. In our assessment, this is an interpretation of the cited results, not a general statement about all techniques.

That tension is, in our assessment, the crux. The cited EDPB documents and analyses acknowledge that AI models can contain personal data and discuss that deletion requests can also extend into model memory. Practice shows, according to the sources, that completely clean forgetting is often only possible through costly retraining, while unlearning remains imperfect. This indicates, in our analysis, that an intention does not suffice; a demonstrable process is needed, in our assessment.

Deletion and retention as a verifiable workflow

On the basis of these sources, we outline, in our analysis, a possible end-to-end workflow that could cover the entire trajectory, from scraping to output. The elements of quarantine and auditable decision logs are our own recommendations here, inspired by but not literally prescribed in the guidelines:

  • Data provenance and data scope. Determine which sources have been scraped and where certain personal data ends up. Without provenance registration, a deletion request cannot, in our assessment, be carried out in a targeted manner.
  • Exclusion and mitigation. Guidelines 03/2026 note, among other things, that controllers can exclude risky sources and should avoid special categories of data as much as possible and, where they do occur, take mitigating measures, including swift removal from training datasets and limiting their appearance in outputs.
  • Limiting memorisation and regurgitation. After training, according to the guidelines, measures are needed to limit the model from returning personal data verbatim or leaking it via privacy attacks.
  • Choice of measure per request. Our recommendation: quarantine, unlearning or full retraining — depending on impact and risk.
  • Auditable decision logs. Our recommendation: record which datasets, indices and models have been adjusted, with timestamps, without copying sensitive content again in the process.

What this means for controlled AI use

For professionals working with confidential information, the question shifts, in our assessment, from 'may I use AI' to 'can I demonstrate how personal data flows through my AI chain and how I reduce it on request'. That is, in our analysis, precisely the point where a verification layer can add value, not as an own model but as a controllable workflow layer.

I am Vera is a privacy-focused AI verification layer, not a chatbot and not an own language model. The Semantic Privacy Shield carries out preprocessing and anonymisation on EU infrastructure before content is offered to the selected AI models; the workflow is designed to pass on only anonymised content, and if a privacy check fails, nothing is passed on. That set-up can help to limit which personal data ends up in external models and gives more insight into how content flows through the chain. Vera does not promise correct or truthful output and cannot rule out hallucinations or risks; it makes verification and processing steps visible, so that the professional final judgement remains with the user.

In our assessment, the new guidelines underline that model memory should be approached as a controlled, documented risk factor. Anyone deploying generative AI would do well, in our analysis, to design deletion and retention policy now as a demonstrable workflow, rather than waiting until a supervisory authority or data subject asks for it.

← All articles