On 8 July 2026, the European Data Protection Board adopted guidelines on web scraping in the context of generative AI. The core message is sober: as soon as personal data is processed during scraping, that falls under the GDPR. Organisations must then pay attention to a legal basis, transparency, data minimisation and the restriction of special categories of data. This shifts the question from are we allowed to input this? to can we demonstrate what happens with which data?
That shift does not stand alone. In the same period, the European Commission published guidelines on transparency obligations for providers and users of AI systems. These guidelines set out transparency obligations under the AI Act and situate them in the period from 2 August 2026 onward: people must be informed when they interact with AI or see AI-generated content.: people must be informed when they interact with AI or see AI-generated content. AI use may therefore not remain hidden within a workflow.
The same logic, also outside Europe
The movement is broader than EU law. According to The Straits Times, Singapore on 20 July 2026 made AI-specific notifications mandatory for companies that use personal data to train generative AI models. Organisations must inform users about which data types and purposes apply and what opt-out options exist. A recurring point: sensitive personal data is difficult to reverse once it ends up in training.
The context comes from IAPP, which describes the underlying PDPC proposals. These concern accountability, legal bases, risk mitigation, transparency and deployment across the entire AI chain. Privacy-sensitive data in generative AI is therefore not an isolated prompt question, but a series of governance decisions from development to deployment.
That privacy is not merely a technical matter is also stated in NIST's AI Risk Management Framework. Whoever deploys generative AI on confidential or personal data must be able to demonstrate appropriate safeguards, data minimisation and risk mitigation. The responsibility remains with the party that deploys the AI.
What organisations must concretely be able to demonstrate
If you place the EDPB guidelines, the EU transparency rules and the Singaporean line side by side, a consistent picture emerges. Responsible AI use with privacy-sensitive information requires three demonstrable things:
- Source selection and data minimisation: which personal data ends up in which AI workflow, and why that and not more?
- Restriction or anonymisation: how is sensitive data restricted or anonymised before a model processes it?
- Transparency and control: is it visible to data subjects that AI was used, and are outputs and decisions traceable per workflow?
The emphasis lies on demonstrability. Regulators do not ask for a statement of intent, but for evidence that you know what is happening and that you retain a grip on the data.
Where a verification layer can help
This is precisely the point where a verification layer becomes practical. Vera is not a chatbot and not its own language model, but a privacy-focused verification layer for professionals working with confidential or high-trust information. Vera can route a task through selected, independent AI models and thereby make verification steps, corrections, disagreements and sources visible for inspection. This does not remove the risk of hallucinations, but it makes control by the user possible.
For the minimisation question, the Semantic Privacy Shield is relevant. It can replace sensitive values in a document with synthetic, session-only equivalents on EU infrastructure, before AI processing takes place. The AI chain analyses the synthetic version; the original values can subsequently be restored locally. The architecture is designed to send only anonymised content onward. If the privacy check fails, the document is not sent onward — the workflow is fail-closed. That is an architecture description, not a promise of flawless anonymisation or full GDPR compliance.
Uploaded PDFs are processed temporarily for the active run and are not stored permanently; metadata may remain in the session history. Within that protected workflow, professionals can view and edit documents via Vera Office, which uses Collabora Online. Vera Office is not a Microsoft Office plug-in and does not edit anything autonomously; the actions remain with the user.
The verifiable record aligns with what the EDPB and the Commission ask for: being able to show which data was used, which AI interaction was made visible and which control and minimisation steps were carried out. The news is not the console, but the shift that lies beneath it.
The common thread
July 2026 marks a turning point: from the EDPB to the PDPC, the same message sounds. Whoever deploys AI on privacy-sensitive information must not only make something technically possible, but demonstrably limit it, inform about it and keep it verifiable. Tools can make those steps visible and support them. The professional final judgement — what you process, what you send onward and what you publish — remains with the user.