The legal framework assumes that protecting privacy means controlling what data is held about a person. AI inference creates a challenge that this framework was not designed for: the ability to reconstruct sensitive information about individuals from data that does not directly contain that information. The reconstructed inference is not the original data. The privacy harm it can cause is equivalent.
The Inference Problem
A model trained on population-level data learns statistical relationships between observable attributes and sensitive characteristics. A model trained on purchasing data learns which purchasing patterns correlate with health conditions, financial stress, or relationship status. A model trained on behavioral data learns which patterns correlate with political affiliation, religious practice, or sexual orientation. The model does not hold data about any individual's health condition, political views, or sexual orientation. It has learned to infer these things from data that does not directly contain them.
The inference is probabilistic, not certain. The model will be wrong about some individuals. For the individuals about whom it is right — which, at scale, is most of them — it has produced sensitive personal information from non-sensitive inputs. The organization that runs the inference has generated sensitive data it did not collect, did not receive consent to hold, and may not be aware it has produced.
This is not a theoretical concern. Inference attacks on machine learning models have demonstrated that models trained on medical records can infer specific individuals' health conditions. Models trained on location data can infer home and work addresses, places of worship, and healthcare providers. Models trained on purchase history can infer pregnancy, chronic illness, and financial difficulty with accuracy sufficient to produce actionable commercial intelligence — which is to say, accuracy sufficient to cause privacy harm.
The privacy framework that governs what data is collected cannot fully govern what a model can reconstruct from that data. The reconstruction capability exists in the model. The harm potential exists in the inference output. The regulatory framework has not fully resolved what obligations these create.
Where Organizations Are Exposed
The Marketing Model and Special Category Inference
AI models used for marketing personalization are trained on behavioral, transactional, and demographic data. These models optimize for predictions that drive commercial outcomes: who is likely to respond to which offer, which customers are at risk of churn, which prospects are in a purchasing mindset. The predictions that drive commercial outcomes frequently correlate with characteristics that are legally protected under privacy regulations — health status, financial circumstance, family situation.
An organization that uses an AI model to target customers who are likely to be experiencing financial difficulty has used inference to identify a sensitive characteristic from non-sensitive inputs. Whether the organization collected financial data about these customers is beside the point. The inference has identified the characteristic. The targeting has acted on it. The GDPR's restrictions on processing sensitive personal data apply to the characteristic regardless of how it was identified.
The HR Model and Protected Characteristic Inference
AI models used in HR processes — screening resumes, scoring interviews, predicting employee performance or attrition — may infer protected characteristics from proxy variables in the training data. A model trained on historical hiring decisions that correlated positively with certain educational institutions, communication styles, or name patterns may have learned to discriminate based on characteristics that correlate with race, gender, or national origin without directly using those protected attributes.
The organization did not input protected characteristic data. The model inferred the correlations from proxy variables in the training data. The discriminatory output is present. The legal exposure under employment discrimination law and data protection law applies regardless of the mechanism by which the discriminatory pattern emerged.
The Health Inference from Consumer Data
Consumer behavior data — purchases, search queries, app usage, location patterns — enables inference of health conditions with commercially significant accuracy. Organizations that hold consumer behavioral data and deploy AI models against that data for commercial purposes may be generating inferences about health conditions that they did not collect, that users did not consent to have generated, and that constitute special category data under GDPR regardless of the mechanism of generation.
The Regulatory Landscape
Regulatory guidance on AI inference and privacy is developing faster than most organizations' legal and governance teams are tracking. The EDPB has published guidance that addresses profiling under GDPR in terms that encompass inference-based processing. Recital 71 of the GDPR explicitly references automated processing that produces profiles about individuals — a description that encompasses the inference outputs of AI models trained on behavioral data.
The EU AI Act's requirements for bias testing in high-risk AI systems implicitly address the inference problem: a system that infers protected characteristics from proxy variables and uses those inferences in consequential decisions is a system whose outputs may reflect discriminatory bias that bias testing is designed to detect. The Act's transparency requirements apply to systems that make or contribute to decisions affecting individuals — which encompasses systems whose inference outputs inform such decisions.
The FTC's approach to AI governance in the United States has addressed inference risk through the lens of unfair and deceptive practices — using AI to infer sensitive characteristics that consumers did not consent to have inferred and using those inferences in ways that cause consumer harm. The regulatory framework is jurisdiction-specific and evolving. The risk is consistent across jurisdictions.
What Governance of Inference Risk Requires
Governing inference risk requires extending the privacy impact assessment methodology to encompass what a deployed AI model can infer, not only what data it directly processes. This requires asking, for each significant AI deployment: what sensitive characteristics could this model infer from its inputs? Under what conditions would those inferences be accurate enough to cause privacy harm? What would the organization do with those inferences, directly or indirectly, in ways that affect individuals?
It requires building model output auditing that looks not only at the accuracy and fairness of model predictions but at what the predictions reveal about individuals who did not consent to that revelation. Proxy discrimination testing — assessing whether model outputs correlate with protected characteristics even when those characteristics were not in the training data — is the technical complement to the legal analysis.
It requires honest conversation between legal, privacy, and data science teams about what the models the organization has deployed are actually doing and what the privacy implications of their outputs are. Those conversations are often technically complex and organizationally uncomfortable. They are also necessary.
The inference is the risk. Build the governance around what the model produces, not only around what data it was given.
Audit what your AI models can infer. The privacy obligation follows the inference, not only the input data. Govern accordingly.
Enterprise practitioner perspective. Not legal advice. Part of the Deep Trust Governance Series by Verisq. Get the free weekly Breach Digest.
