The most popular advice about GDPR privacy by design is also the least useful: “build privacy in from the start.” Article 25 requires more than an early policy review or a privacy statement attached to a product launch. It requires controllers to make defensible technical and organisational choices, preserve evidence of those choices, and revisit them while processing continues. For an engineering or governance team, the practical question is therefore not whether a system is “privacy friendly.” It's whether someone can inspect the architecture, defaults, supplier arrangements, access rules, risk assessment, and operational logs, then understand why the system processes personal data in that way.
What Privacy by Design Actually Means Under Article 25
Article 25 makes privacy by design a lifecycle engineering duty, with accountability distributed across the people who build, supply, configure, and operate a system. Under Article 25(1), controllers must implement appropriate technical and organisational measures when determining the means of processing and throughout the processing itself. The EDPB Guidelines 4/2019 identify measures such as data minimisation, pseudonymisation, and encryption as controls to incorporate into architecture, not additions after deployment. Pseudonymised information remains personal data when additional information could enable re-identification, so residual risk must remain visible in the design.
A privacy statement records intent. It does not demonstrate implementation. Inspectable evidence includes requirements records, data-flow diagrams, supplier instructions, configuration decisions, test results, retention rules, access reviews, and risk assessments. Builders can document the control, vendors can expose relevant settings and processing details, and deployment teams can preserve evidence that production configuration matches the approved design. The controller remains accountable for the processing, but accountability depends on artefacts distributed across that chain.
Design and default are separate duties
Article 25(1) concerns safeguards embedded in processing operations. Article 25(2) concerns what happens by default. It requires the personal data necessary for each specific purpose to be processed by default, limiting the amount collected, processing scope, storage period, and accessibility. Personal data should not be accessible without the individual's intervention to an indefinite number of people. Article 25(2) GDPR
A system can use strong encryption and still fail the default obligation if optional fields are enabled, retention lacks a defined purpose, or records reach more people than necessary. A restrictive interface cannot correct an architecture that sends clear identifiers to an external service. These trade-offs should appear in configuration baselines, deployment checks, and exception approvals.
For sensitive support workflows, privacy by design for customer support shows how these duties translate into operational choices. In AI-enabled work, the same lifecycle logic applies to models, prompts, integrations, and output paths, as outlined in GDPR responsibilities in a generative AI workflow.
Practical rule: Treat every privacy claim as a control that should produce inspectable evidence.
The Four Factors Controllers Must Balance
Article 25 does not demand one architecture for every controller. It demands a documented choice based on four factors: the state of the art, the cost of implementation, the nature, scope, context and purposes of processing, and the risks to people's rights and freedoms. The ICO presents these factors as the basis for selecting appropriate design measures. ICO guidance on data protection by design and by default
Read the factors as one assessment
State of the art means identifying technical and organisational measures reasonably available for the processing, rather than accepting a vendor's newest feature. Encryption, tokenisation, segmentation, access restrictions, privacy testing, and controlled deletion may all be relevant. The design record should explain which safeguards were selected, why they fit the processing, and what builders, suppliers, or deployment teams must implement.
Cost supports proportionality, not a general exemption. The controller should record implementation effort, operational constraints, alternatives considered, and the reason for the chosen control. A limited budget does not make a foreseeable high-impact risk acceptable. The resulting decision should be visible in approval records, procurement documents, or an exception log.
Nature, scope, context and purposes establish what problem the controls must address. A confidential legal workflow, an employee-monitoring system, and a public-facing contact form involve different data, affected people, recipients, purposes, and operating conditions. Those details should appear in the processing description, system design, data-flow map, and supplier configuration.
Risks to rights and freedoms need more than a generic threat label. The analysis should describe possible harm to identifiable people, assess likelihood and severity, and link each risk to a safeguard. It should also record who accepted any residual risk and when that decision must be revisited.
A defensible Article 25 file therefore contains more than four headings. It connects each factor to a decision, an accountable owner, an inspectable implementation artefact, and a review trigger. That structure makes responsibility testable across the controller, its builders, vendors, and deployment teams.
Data Minimisation and Privacy by Default in Practice
Article 25(2) becomes useful when a team translates it into four control surfaces: collection, processing extent, retention, and accessibility. These are not abstract privacy principles. They're places where product requirements, code, configuration, and operational review can either restrict or expand exposure.
The four control surfaces
At collection, a form or ingestion interface should have a field-level necessity rationale. If a field has no defined purpose, the default should be to omit it, not collect it “in case it becomes useful.” This doesn't prevent legitimate collection. It forces the team to connect each field to a specific processing purpose.
During processing, rules should remove or suppress attributes that the next operation doesn't need. A summarisation service may need document text but not direct identifiers. An analytics pipeline may need an event category without retaining the underlying content. The control is inspectable when the transformation rule, test case, and resulting payload can be reviewed.
Retention requires an enforceable schedule, not just a statement that data won't be kept longer than necessary. Deletion or archival logic should identify the purpose, the event that starts the retention period, the responsible system, and the evidence that the rule ran. Where longer retention serves a separate purpose, that purpose should be assessed rather than inherited from the original workflow.
Accessibility concerns both internal users and onward dissemination. Role-based scopes, tenant isolation, approval gates, and restricted exports can prevent personal data from becoming available to an indefinite audience. The ICO's guidance also stresses that pseudonymisation, encryption, and other measures should be selected in light of the processing context and risk, not treated as substitutes for minimisation. ICO pseudonymisation guidance
- Amount collected — Inspectable Control: Field-level necessity decisions, constrained forms, and ingestion filters
- Extent of processing — Inspectable Control: Payload rules, attribute suppression, and purpose-specific processing paths
- Storage period — Inspectable Control: Enforced retention schedules, deletion jobs, and execution records
- Accessibility — Inspectable Control: Role scopes, export restrictions, approval gates, and access logs
The trade-off is usually visible in analytics. Keeping more fields for longer may support richer analysis, but it also expands the information available for later use, access, and disclosure. A privacy-by-default design makes that trade-off explicit and requires an active decision before expanding collection or visibility.
For AI workflows, data minimisation in a generative AI workflow provides a useful lens for examining what reaches a model, what remains in logs, and what the user can export.
Pseudonymisation as a Layered Design Control
Pseudonymisation is a risk-reduction architecture, not a conversion of personal data into anonymous data. Its protection depends on controlling the information that permits reversal. Under Article 25, pseudonymisation is an example of an appropriate measure, while pseudonymised data remains personal data when separately held information can support re-identification.
Four layers need four kinds of evidence
Identifier replacement substitutes direct identifiers with tokens or other pseudonyms. The implementation record should specify what is replaced, how collisions are handled, and whether the transformation preserves linkability across events.
Isolated mapping data separates reversal information from the operational dataset. The ICO recommends distinct physical or logical storage, with network segmentation where appropriate, rather than placing the mapping beside pseudonymised records.
Access controls on re-identification establish a separate permission boundary. Staff or services able to use the mapping should have restricted, authenticated access, with monitoring independent of ordinary operational-data access.
Re-identification limits require documented tests and decisions. A DPIA or equivalent processing record should state when reversal is allowed, who authorises it, what logging is required, and how linkage risk is reassessed after system changes.
Encryption protects mapping material, but it does not determine who may decrypt it or under which conditions. Hashing can reduce direct exposure while still permitting linkage or inference, depending on the implementation and available auxiliary information.
A defensible design preserves four inspectable artefacts: replacement logic, storage topology, access reviews, and re-identification test results. The organisation should also record whether pseudonyms are stable, rotated, purpose-scoped, or shared across systems. Those records distribute accountability across the teams that build the control, vendors that operate components, and deployment teams that configure access and exceptions.
Residual risk matters: Pseudonymisation reduces linkage risk. It does not remove the controller's responsibility to assess whether people can still be identified.
The distinction between pseudonymisation and anonymisation, particularly in AI workflows, is developed further in pseudonymisation versus anonymisation in AI workflows.
The following visual demonstrates the layered model in operation:
Mapping Article 25 to a Confidential Document Workflow
Consider a confidential legal-document workflow that ingests contracts, sends selected context to a large language model for summarisation, and returns findings to counsel. The privacy question isn't whether the model provider offers security features. The team must decide what leaves the customer environment, which identifiers are necessary, how the system blocks accidental disclosure, and what proves those controls operated.
A control-by-control mapping
A semantic privacy layer can detect sensitive entity types such as names, addresses, financial identifiers, and contract clauses, then hash or replace them before a model call. That implements a minimisation and pseudonymisation decision at the outbound boundary. The workflow still needs to document the detection rules, known limitations, and treatment of content that doesn't match a configured entity type.
Local pseudonymisation keeps the mapping on the customer side rather than sending clear identifiers to the model provider. The mapping store should have its own access controls, encryption, operational ownership, and re-identification policy. Without those boundaries, replacing a name in the prompt may only move the risk into another system.
A fail-closed outbound check can block a prompt or response when unredacted personal data is detected. Testing the policy with synthetic adversarial inputs creates evidence that the gate was exercised against deliberate failures, rather than merely enabled in configuration.
What an investigator could inspect
For this workflow, useful artifacts include:
- Redaction records: Which entity classes were detected, transformed, or left unresolved.
- Outbound policy reports: Whether the prompt passed, failed, or required intervention.
- Model prompt captures: The exact protected context sent for processing, subject to appropriate access restrictions.
- Restoration records: When and how local pseudonyms were mapped back into the customer document.
- Configuration snapshots: The policy version, model route, and deployment settings active at the time.
The four Article 25 factors shape the design. Legal text creates a context where confidentiality and individual impact can be significant. The state of the art informs available filtering and isolation measures. Cost matters when selecting controls, but it should be weighed against the risk and the feasibility of alternatives. The result is not a claim of automatic compliance. It's a traceable chain from risk assessment to safeguard to observed operation.
Those features can support evidence collection, but they don't replace the controller's legal assessment, supplier due diligence, or professional judgment.
Distributed Accountability Across Builders, Vendors and AI
Article 25 is often described as a controller obligation, but modern processing rarely sits inside one organisation's architecture. Software builders create the application, SaaS vendors operate components, model providers handle inference or logging, and deployment teams configure integrations. The controller still has responsibility for its processing decisions, yet it may not control every technical layer that affects the result.
Where the chain breaks
A vendor may offer privacy settings without making restrictive defaults configurable or persistent. An AI provider may retain prompts or outputs in operational logs under terms the deployment team hasn't examined. An integration layer may pass complete records to a downstream service because no local filtering occurs before the API call.
Each gap can undermine Article 25(2). A controller may believe that data minimisation is in place, while the connector sends extra fields. A product team may configure restricted visibility, while a vendor update resets the setting. A contract may allocate responsibility on paper, while no one exports the configuration or checks the audit trail.
- Minimised inputs — Controller Responsibility: Define necessary data and approve payload rules; Vendor / Builder Responsibility: Provide filtering and stable configuration controls; AI Provider Responsibility: Process only agreed inputs and document handling
- Protective defaults — Controller Responsibility: Set purpose-specific defaults and test them; Vendor / Builder Responsibility: Preserve restrictive settings through updates; AI Provider Responsibility: Offer clear controls for retention, logging, and access
- Pseudonymisation — Controller Responsibility: Decide when reversal is permitted and govern the mapping; Vendor / Builder Responsibility: Support tokenisation, isolation, and access boundaries; AI Provider Responsibility: Avoid receiving clear identifiers where they aren't needed
- Evidence and review — Controller Responsibility: Maintain DPIA, decisions, and verification records; Vendor / Builder Responsibility: Supply change notices, logs, and configuration exports; AI Provider Responsibility: Provide relevant processing and operational evidence
A practical control matrix should assign each safeguard to an owner and name the evidence that travels with the deployment. That evidence might include exported policies, configuration snapshots, supplier terms, access logs, and test results.
Teams building broader governance programs can compare this supply-chain approach with a compliance framework for Canadian SMBs, while keeping the legal analysis jurisdiction-specific. Canadian governance material can inform process design, but it doesn't replace the GDPR obligations that apply to EU processing.
The contrarian point: In a shared architecture, documentation isn't administrative residue. It's how the controller demonstrates which party made each privacy-relevant decision.
Why Design Failures Carry a Defined Fine Range
Article 25 is not a soft design preference. A failure can fall within a defined GDPR enforcement tier. The EDPB's fine-calculation guidance places infringements of data protection by design and by default within the lower administrative-fine category, up to €10 million or 2% of worldwide annual turnover, whichever is higher. The higher tier reaches up to €20 million or 4% of worldwide annual turnover, whichever is higher, for obligations covered by that tier. EDPB fine-calculation guidelines
Those figures are not a pricing model for non-compliance. They define the possible exposure, while the design record helps explain how an authority may assess the failure. After an incident or complaint, investigators can examine what the controller knew about the processing, which safeguards were available, what defaults shipped, how suppliers were configured, and whether the organisation revisited its risk assessment.
Why the table cannot honestly name cases here
A comparison of supervisory-authority cases requires case-specific primary decisions. The verified material establishes the fine bands and the classification of Article 25 infringements, but it does not provide verified case names, authority-specific decisions, or documented design failures for Germany, France, Ireland, or the Netherlands. Filling those cells with plausible-sounding examples would turn an evidence table into speculation.
- EU GDPR framework — Maximum Fine Tier 4: Up to €10 million or 2% of worldwide annual turnover, whichever is higher; Notable Article 25 Case: Case-specific primary decision requires confirmation; Design Failure Cited: Case-specific finding requires confirmation
- National supervisory authorities — Maximum Fine Tier 4: Applied within the GDPR fine framework; Notable Article 25 Case: Case-specific primary decision requires confirmation; Design Failure Cited: Case-specific finding requires confirmation
The useful operational conclusion is that inspectable design evidence reduces uncertainty during review. A controller that can show the processing purpose, the four-factor assessment, default settings, pseudonymisation design, supplier allocation, test results, and later reassessments can demonstrate an active control process. That evidence also distributes accountability: builders can substantiate implementation, vendors can substantiate their configured service, and deployment teams can show that restrictive settings survived release. A policy statement alone gives the controller less basis for explaining why the shipped system was appropriate.
The ICO's current guidance reflects the changing scope of privacy-by-design work. Its update adds a subsection on children's higher-protection duties under UK GDPR after the Data (Use and Access) Act 2025. The change illustrates why design obligations require lifecycle monitoring, especially where sector-specific duties affect defaults, testing, or review triggers.
More articles on this subject are collected in the AI privacy and GDPR overview.
Sources and references
Sources: Core legal claims are attributed to Article 25 GDPR (gdpr-info.eu), the EDPB Guidelines 4/2019 and ICO guidance on data protection by design, pseudonymisation and fine calculation, with design mappings presented as editorial analysis.