It is a mature discipline with established practices: model inventory, validation, ongoing performance monitoring, periodic revalidation, and governance structures that include model risk committees with authority over model deployment. Organizations applying this framework to generative AI models are discovering that it was designed for a different kind of model — one with bounded outputs, defined performance metrics, and stable behavior in production. Generative AI models have none of these properties.
What Traditional MRM Was Built For
The quantitative models that traditional model risk management was designed to govern share specific characteristics. They take defined inputs and produce bounded outputs — a credit score, a risk exposure estimate, a trading signal. Their performance can be measured against ground truth — the loan either defaulted or it did not. Their behavior is deterministic or near-deterministic — the same inputs produce the same or predictable outputs. They are validated through backtesting against historical data. And they are stable once deployed — their parameters do not change unless the model is retrained.
Generative AI models share almost none of these characteristics. They take natural language inputs of virtually unlimited variety and produce natural language outputs with no upper bound on their content. Their 'performance' is multidimensional and often subjective — accuracy, helpfulness, safety, tone, and factual correctness are all dimensions of performance that cannot be collapsed into a single metric. Their behavior is probabilistic — identical inputs may produce different outputs. Backtesting against historical data produces accuracy measures that do not translate to the full range of production use cases. And they change continuously through vendor updates.
Traditional model risk management assumes that models can be fully characterized before deployment and that their behavior in production will be consistent with their validated behavior. Generative AI models cannot be fully characterized before deployment and their production behavior includes edge cases that no validation exercise will have assessed.
The Gaps When Traditional MRM Meets Generative AI
Validation That Cannot Cover the Output Space
Traditional model validation tests model performance against defined scenarios and measures outcomes against ground truth. Generative AI model validation faces a fundamental scope problem: the output space of a generative model is effectively unbounded. A model that generates text can produce outputs on any topic, in any style, with any degree of accuracy or inaccuracy, appropriate or inappropriate content, factually correct or hallucinated assertions. No validation exercise can test more than a fraction of this space.
The validation approach that has emerged for generative AI — red teaming, adversarial testing, evaluation against curated benchmark datasets — provides evidence about model behavior in the tested scenarios. It does not provide evidence about model behavior in the full production scenario space. The gap between tested scenarios and production scenarios is where the validation did not look.
Performance Monitoring Without a Clear Performance Metric
Traditional model monitoring tracks performance against a defined metric: credit model accuracy against default rates, trading model performance against P&L. For generative AI models used in enterprise workflows, the performance metric is often ambiguous. Is the customer support AI performing well? How is performance defined? Response accuracy? User satisfaction? Escalation rate? Harmful content rate? Each metric captures something, and none captures everything. Performance monitoring for generative AI requires a portfolio of metrics, which requires a portfolio of monitoring infrastructure.
The Foundation Model Problem
Most enterprise generative AI deployments use foundation models developed by AI vendors accessed through APIs. The model risk management framework that governs an internally developed quantitative model does not map directly to a foundation model whose internal architecture is proprietary, whose training data is not disclosed in full, whose updates are made by the vendor without customer approval, and whose behavior changes with each vendor update. The traditional MRM workflow assumes the organization can inspect the model. Foundation models are not inspectable.
What an Adapted MRM Framework Requires
Adapting model risk management for generative AI requires specific additions to the traditional framework rather than replacement of it. The inventory and lifecycle governance that traditional MRM provides is still needed. What is added:
Tiered validation proportionate to deployment risk. High-stakes deployments — AI used in consequential decisions about individuals — require more rigorous validation including red teaming and adversarial testing. Lower-stakes deployments — AI used for internal productivity — require less rigorous validation. Risk-proportionate validation allocates the validation investment where the deployment risk warrants it.
Multi-dimensional monitoring dashboards. Monitoring that tracks accuracy, harmful content rates, hallucination rates, user feedback signals, and escalation rates simultaneously produces a performance picture that single-metric monitoring cannot. The dashboard design requires deciding which dimensions matter most for each deployment context.
Vendor update governance. Foundation model updates by vendors should be treated as model changes that require governance review. Behavioral testing after vendor updates, compared to pre-update baselines, detects changes that affect the deployment's performance or safety characteristics.
Use case boundary documentation. Generative AI models deployed for a defined use case should have documented boundaries: the scenarios the model was validated for and the scenarios it was not validated for. These boundaries govern where the model can be deployed and where deployment requires additional validation.
Build the adapted framework. Traditional MRM is the foundation. Generative AI requires an additional floor.
Take what traditional MRM provides. Add the validation, monitoring, and vendor governance that generative AI specifically requires. Neither alone is sufficient.
Enterprise practitioner perspective. Not legal advice. Part of the Deep Trust Governance Series by Verisq. Get the free weekly Breach Digest.
