The deletion assumption, the purpose limitation assumption, and the data minimization assumption all behave differently once data has become training input.
Why This Matters Now
Privacy law was developed in an era when data existed as records. The foundational rights of privacy law, including access, rectification, erasure, and portability, reflect an assumption that personal data is a discrete, locatable artifact that can be retrieved, modified, or removed.
AI training transforms that artifact. When personal data is used to train a model, the data is processed and the model learns from it. The data is then no longer needed in its original form. But what the model learned from the data persists in its weights, influencing every output the model produces, in ways that have no mapping to any of the foundational privacy rights.
Privacy law governs data. AI training creates knowledge derived from data. These are not the same thing. The privacy frameworks that govern the first do not adequately govern the second.
The Governance Problem Beneath the Surface
The governance problem is not that privacy law is inapplicable to AI. It applies. The problem is that the assumptions embedded in privacy law stop functioning as designed when applied to AI training scenarios.
Take the right to erasure. It assumes the data exists in a form that can be located and deleted. After AI training, the data has been transformed into model weights. Erasure of the original data does not affect the model weights. Take purpose limitation. After AI training, the model can be used for purposes not defined at collection. The limitation that applies to data does not apply to the model's capabilities.
What This Actually Means in Enterprise Practice
Deletion of Training Data Does Not Delete Model Knowledge
When a data subject exercises a right to erasure, the privacy team deletes the individual's records from all data stores. If that individual's data was used in AI training before the request, the model's weights still encode patterns learned from that data. The data is gone. The knowledge persists.
Erasure requests fulfilled for training data without addressing model knowledge are incomplete erasure fulfillments. The privacy program records completion. The AI system continues to operate on learned patterns from deleted data.
Lawful Basis for Collection Does Not Authorize All Training Use
Data collected under one lawful basis for a specific purpose may be used for AI training without adequate assessment of whether the training use falls within the original lawful basis. Consent for service personalization does not automatically authorize use of the same data for training a general-purpose language model.
Every AI training dataset should be assessed against the lawful basis under which each data element was collected. In most organizations, this assessment has not been performed for existing training datasets.
Anonymization Before Training Does Not Guarantee Post-Training Anonymity
Anonymizing data before using it for AI training reduces but does not eliminate privacy risk. Models trained on anonymized data may develop inference capabilities that allow re-identification of individuals from anonymized inputs. Membership inference attacks can determine whether specific individuals' data was in the training set.
Purpose Limitation Is Structurally Incompatible with Model Training
A model trained on data collected for purpose A acquires capabilities that extend to purposes B, C, and D. The model's capabilities are not limited by the collection purpose of its training data. This creates a structural incompatibility between GDPR purpose limitation requirements and model training, which produces capabilities that extend beyond any specified purpose.
How Different Teams See This: Where They All Miss
AI training privacy governance requires technical and legal expertise that does not exist fully in any single team. The gap exists at the intersection of privacy law, statistical learning theory, and governance program design.
Framework Control Reference
The specific control obligations most relevant to this topic. Use in governance discussions, vendor assessments, and audit responses.
These controls share a common requirement: the obligation is active, not declarative. Documenting alignment is not the same as demonstrating it.
The Enterprise Reality Gap
The enterprise AI training data privacy reality gap is between the privacy governance coverage organizations document for AI training and the actual privacy risk profile of trained models. The documentation covers the data layer. The risks at the model layer are not captured in standard DPIA methodology.
Privacy governance for AI requires understanding both what the privacy framework says and where it stops functioning as designed. The gap between those two things is where the most significant AI privacy risks currently operate.
Enterprise Scenario
The anonymization was correctly applied to the training data. The model learned from that data in ways that preserved more individual-level signal than the anonymization was designed to remove. The DPIA assessed the data layer risk. It did not assess the model layer risk. Both existed. Only one was governed.
Industry Signal
The EDPB's Opinion 28/2024 on AI models and personal data is the most significant regulatory signal on this topic to date. The opinion addresses whether trained AI models themselves constitute personal data and what obligations this creates. The direction of the opinion signals that regulators intend to extend privacy obligations to trained models, not just training datasets.
The regulatory question is moving from whether privacy law applies to AI training to how it applies to the model that training produces. The second question is harder, and the governance infrastructure required to answer it does not yet widely exist in enterprise privacy programs.
Enabling Capabilities
- Privacy-preserving machine learning: Differential privacy, federated learning, and secure aggregation techniques that reduce the privacy risk of AI training without eliminating model capability.
- Training data lineage with privacy mapping: Systems that connect personal data records to training datasets with lawful basis documentation for training use.
- Membership inference testing: Technical assessment of whether deployed models retain individual-level signal sufficient to enable re-identification.
- AI-specific DPIA methodology: Extended DPIA frameworks that address model-layer privacy risks alongside data-layer risks.
A Practical Starting Point
For your highest-risk AI training datasets, perform a model-layer privacy assessment alongside the standard DPIA. Specifically: what inference capabilities has the trained model acquired beyond the data attributes in the training set? Can the model be used to re-identify individuals from apparently anonymized inputs? Has a membership inference assessment been conducted?
The DPIA covers the governance the standard methodology enables. The model-layer assessment covers the governance that AI training specifically requires. Both are needed.
Questions Leaders Should Be Asking
- For our AI training datasets, have we assessed whether the trained model retains individual-level signal that enables re-identification beyond the anonymization applied to training data?
- Have we assessed the lawful basis for each training dataset specifically for the training use case, or have we assumed that the original collection basis extends to training?
- What is our governance position on erasure obligations for individuals whose data trained AI models, and has that position been reviewed by legal, privacy, and AI governance together?
- Are differential privacy or other privacy-preserving ML techniques used in our highest-risk AI training programs, and if not, has the decision not to use them been assessed against the privacy risk it accepts?
What to Require From Vendors
Ask directly:
"For AI systems trained on data we provide, what privacy-preserving training techniques are used, has the trained model been assessed for membership inference vulnerability, and what documentation supports compliance with GDPR training data obligations?"
Expect as evidence:
- Documentation of privacy-preserving training techniques used
- Membership inference assessment results or vendor position on assessment
- Training data lawful basis documentation with specific analysis of training use scope
- Model-layer privacy risk assessment or documentation of residual risk position
A vendor who confirms DPIA completion for training data without addressing model-layer privacy risks has governed the data and left the model ungoverned.
Demonstrating Diligence
- Documentation: AI training DPIA with model-layer risk assessment; training data lawful basis documentation with training use analysis; membership inference assessment records.
- Process: Joint privacy and AI governance review for AI training programs; model-layer privacy assessment as a required step in the AI development lifecycle.
- Technical evidence: Privacy-preserving technique implementation records; membership inference assessment outputs.
AI training privacy diligence requires showing governance at the model layer, not just the data layer.
Closing Perspective
AI training data breaks the assumptions that privacy governance was built on. That is not a criticism of privacy governance. It is an accurate description of what happens when a governance framework encounters technology that operates beyond its design assumptions.
The organizations that navigate this most effectively are those that recognize the limitation honestly and build governance specifically for the model layer that standard privacy frameworks do not adequately address.
Privacy law governs data. AI training creates knowledge. Govern both.
Enterprise practitioner perspective. Not legal advice. Part of the Deep Trust Governance Series by Verisq. Get the free weekly Breach Digest.
