These claims describe the model's safety in the test conditions and populations the vendor evaluated it against. They do not describe the model's safety for the specific deployment context, the specific user population, the specific data being processed, and the specific consequential decisions the model will influence in the specific organization's use case. Safe in general and safe for your specific deployment are different claims.
What Vendor Safety Claims Actually Cover
AI vendor safety evaluations are conducted against the threat models and use cases the vendor considers most significant. Red teaming exercises test the model against adversarial inputs that the red team considered most likely to produce harmful outputs. Content safety evaluations measure harmful output rates across a distribution of inputs that the vendor compiled. Bias evaluations test the model's performance across demographic groups represented in the vendor's evaluation datasets.
These are legitimate and valuable evaluations. They produce evidence about the model's behavior in the tested conditions. The limitation is that the tested conditions are the vendor's conditions — not the organization's specific deployment context, user population, data characteristics, or use case. A model evaluated for safety in a general consumer context may behave differently when deployed in a healthcare context with clinical notes as inputs. A model evaluated for bias across the demographic distribution in the vendor's dataset may exhibit different bias patterns in the organization's specific customer population.
The vendor's safety claim is an assertion about the model's general safety properties. The organization's use case may expose safety risks that the general evaluation did not assess.
Vendor safety evaluation answers: is this model safe under the conditions we tested? Deployment safety assessment answers: is this model safe for our specific use case, our specific users, and our specific data? The first question must be answered before the second. Both questions must be answered before deployment.
The Gaps Between General Safety and Deployment Safety
Domain-Specific Risk That General Evaluation Does Not Capture
A model deployed in financial services to generate investment analysis may produce outputs that are not harmful in a general context but are problematic in the financial context: statements that could be construed as investment advice without the required disclosures, analysis that systematically favors certain asset classes due to training data composition, or explanations that a non-specialist user might misinterpret in consequential ways. General content safety evaluation does not assess domain-specific risks. Domain-specific assessment requires domain expertise that the vendor's evaluation team may not have.
User Population Characteristics Not Reflected in Evaluation
A model evaluated for safety against an adult general population may not have been evaluated for the specific risks that emerge when deployed to vulnerable populations: users with mental health conditions, users with limited technical literacy who may not understand the model's limitations, or users in contexts where the stakes of model outputs are higher than in the general consumer context. The organization that deploys the model to a population with specific vulnerability characteristics should assess whether the vendor's safety evaluation covers those characteristics.
Data Sensitivity That Changes the Risk Profile
A model that is safe when processing public information may create safety risks when processing sensitive personal data, confidential business information, or legally privileged content. The safety risks are not only about harmful outputs — they include risks of data exposure through model outputs, risks of the model producing outputs that reveal information the user should not have access to, and risks of the model retaining sensitive information in ways that create privacy obligations. General safety evaluation does not assess data-sensitivity-specific risks.
Conducting Deployment-Specific Safety Assessment
Deployment-specific safety assessment requires extending the vendor's general evaluation with organization-specific testing. The extension has three components.
Domain-specific red teaming. Testing the model against adversarial inputs specific to the deployment domain — the financial manipulation scenarios, the medical misinformation scenarios, or the legal liability scenarios that the domain presents. This requires domain expertise in the red team.
User population representation. Evaluating the model's outputs for the specific user population in the deployment, including any vulnerability characteristics of that population that the vendor's evaluation may not have covered.
Data sensitivity testing. Testing whether the model behaves safely when processing the specific data categories present in the deployment — not only the data categories the vendor evaluated against.
The deployment-specific assessment supplements rather than replaces the vendor's evaluation. The vendor's evaluation provides the baseline. The organization's assessment determines whether the baseline is sufficient for the specific deployment.
The vendor evaluated the model in their context. Evaluate it in yours before deployment.
Conduct deployment-specific safety assessment before deploying any consequential AI system. The vendor's safety claim covers their context. Your context requires your assessment.
Enterprise practitioner perspective. Not legal advice. Part of the Deep Trust Governance Series by Verisq. Get the free weekly Breach Digest.
