The AI Learning Hub Journal

The Model Risk Frame You Already Have

The frame you already have — and the five places it strainsmodel risk was defined long before this technology, and these features land inside it TWO SOURCES OF MODEL RISK 1 · the model is wrongincorrect output drives the decision2 · it is sound, used outside its purposethe more common and more expensive ofthe two, and the harder one to see WHAT VALIDATION COVERSconceptual soundness — design and assumptions fitthe intended useongoing monitoring — performance tracked againstthresholds after deploymentoutcomes analysis — results against realisedoutcomes and sensible benchmarks MODEL RISK MANAGEMENT DEVELOPMENT AND USErobust design, documentedassumptions, and use withinthe stated purpose INDEPENDENT VALIDATIONa separate exercise, bypeople with the standingand incentive to say no GOVERNANCE AND CONTROLSpolicies, committees andthe authority that makesthe first two happen US banking supervision and the UK prudential regulator both set these expectations; other regimes mirror them EFFECTIVE CHALLENGE — MOSTLY AN ORGANISATIONAL PROPERTY validation is not testing by the build team — it needs competence, standing and the incentive to say noit fails quietly when the validator reports to the model owner, is under-resourced,or arrives after the launch date has already been announced WHERE THE FRAME STRAINS — ADAPT THE TECHNIQUES, NEVER SKIP THE VALIDATIONnothing to backtest — nosingle realised outcome fora text outputthe prompt and the retrievalcorpus are part of the modeland belong in its documentsidentical inputs can producedifferent outputs, breakingreproducibility teststhe population the systemsees has no stabledefinitionthe provider controls theweights — it can change withno release on your side BRING IT INTO SCOPE AND TIER IT BY MATERIALITY arguing that it is not a model is the answer that fails — and misuse outside intended purpose is the costlier source
Bring an AI feature into the model risk regime and tier it by materiality — adapt the techniques, never skip the validation.

Model Risk Was Defined Long Before AI

Model risk is conventionally defined as the potential for adverse consequences from decisions based on incorrect or misused model output, and it has two sources: the model may be fundamentally wrong, or it may be sound but used outside the purpose it was built for. The second source causes more damage in practice than the first. Supervisory guidance in US banking sets expectations across three areas — robust development, implementation and use; effective independent validation; and governance, policies and controls that make the first two happen. The UK prudential regulator has published its own model risk management principles for banks, and comparable expectations appear in other jurisdictions. This frame is older than the current technology and it is where AI features land.

  • Two sources of model risk: the model is wrong, or it is right and used outside its intended purpose
  • Misuse outside intended purpose is the more common and more expensive failure of the two
  • The frame has three parts: development and use, independent validation, and governance and controls
  • US banking supervision and the UK prudential regulator both set model risk expectations; other regimes mirror them

Independent Validation and Effective Challenge

Validation is not testing performed by the team that built the model. It is a separate exercise carried out by people with the competence, the standing and the incentives to say no — the quality usually described as effective challenge. It covers conceptual soundness, meaning the design and assumptions make sense for the intended use; ongoing monitoring, meaning performance is tracked against thresholds after deployment; and outcomes analysis, meaning results are compared against realised outcomes and sensible benchmarks. Challenge fails quietly when the validator reports to the model owner, is under-resourced, or arrives after the launch date has been announced publicly — which is to say that most validation failures are organisational rather than technical.

  • Validation is independent of development, with authority and standing to reject the model
  • Three components: conceptual soundness, ongoing monitoring, and outcomes analysis against benchmarks
  • Effective challenge requires competence and incentive — one without the other produces a rubber stamp
  • A launch date announced before validation completes has already decided the validation outcome

Where AI Features Strain the Frame

The frame holds, but several of its assumptions break. There is often no single realised outcome to backtest a text output against. The input is a prompt and a retrieved corpus, both of which are part of the model whether or not anyone documented them that way. Outputs may vary between identical runs. The population the system sees has no stable definition. And the provider controls the weights, so the model can change without a release on your side. Firms also argue about whether an LLM feature is a model at all under their own policy. The answer that survives supervision is to bring it into scope explicitly, tier it by materiality, and adapt the validation techniques rather than skip the validation. Does Your AI Actually Work? covers how those evaluation sets get built.

  • No single realised outcome to backtest against, and no stable population definition
  • The prompt and the retrieval corpus are part of the model and belong in its documentation
  • Identical inputs may produce different outputs, which breaks conventional reproducibility tests
  • Bring it into scope and tier it by materiality — arguing it is not a model is the answer that fails

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.