Models Are Not New Here
Finance Has Been a Model Industry for Decades
Almost every other sector is meeting statistical decisioning for the first time. Finance is not. Statistical credit scorecards have driven consumer lending decisions since the middle of the last century. Fraud detection moved from rules to learned models decades ago. Pricing and valuation models, rating models, capital and stress models, actuarial models, anti-money-laundering transaction monitoring and algorithmic execution all predate the current wave by many years. Just as importantly, a supervisory apparatus grew up around them: definitions of model risk, expectations for independent validation, inventories, documentation standards and governance committees. When someone in a bank says "we already have a process for this", they are usually right.
- Consumer credit has been scored statistically for generations — automated decisioning is not new territory
- Fraud detection, valuation, capital, actuarial and execution models all long predate language models
- Supervisors built model-risk expectations around those models: validation, inventory, documentation, governance
- The institutional muscle for governing models already exists — the question is whether it stretches to this
What Genuinely Changed
Three things are actually new. First, unstructured data became usable: contracts, filings, emails, call transcripts, policy documents and scanned correspondence were previously either ignored or hand-processed at cost. Second, language became both an input and an interface — a non-technical user can now direct a system in prose, which changes who can invoke a model and how easily. Third, the output is generative: instead of a score or a class label, the system produces text that reads as reasoned. That last change is the disruptive one for governance, because a score can be backtested against an outcome and a paragraph cannot be, at least not the same way.
- Unstructured text — contracts, filings, calls, correspondence — became tractable at volume for the first time
- Language as an interface widens who can invoke a model, often outside the controls built for model users
- Generative output replaces a score with prose that reads as reasoned, whether or not reasoning occurred
- A score has an outcome to backtest against; a paragraph does not, and that is the governance problem
What Did Not Change, and What Was Added On Top
Existing obligations are largely technology-neutral, and none of them were suspended. Nothing about a new architecture removes the duty to validate a model independently before it is relied upon, and nothing removes fair lending duties, anti-money-laundering obligations, records and communications retention requirements, data protection duties, or the requirement that a firm be able to explain a decision it made about a customer. Accountability does not move either: the named individual who owned the outcome before still owns it. What has changed is that some jurisdictions have layered technology-specific rules on top of that baseline, the EU AI Act being the clearest example, and several supervisors have issued expectations addressed to AI directly. Both layers apply at once.
- Validation, documentation and governance duties attach to the use, not to the architecture
- Fair lending, AML and KYC obligations, records rules and data protection were never suspended
- Accountability stays with the named human owner — module 2 takes that into third-party arrangements
- Where an AI-specific regime exists, such as the EU AI Act, it sits on top of the baseline rather than replacing it
Why the Continuity Is Both Comfort and Trap
The comfort is real: you have an inventory, a validation function, a change process and a committee that already knows how to say no. Most institutions do not need to invent a governance regime, only to extend one. The trap is that the existing machinery quietly assumes a particular kind of model — deterministic, with stable numeric inputs, a measurable outcome to backtest, coefficients someone can inspect, and a release you control. A language-model feature can breach every one of those assumptions at once. Teams that force it through unchanged produce validation reports full of metrics that do not mean anything here; teams that declare it out of scope produce no oversight at all. The work sits in deciding which of the old assumptions still hold for this use, and recording that decision.
- Extend the model-risk regime you have rather than building a parallel one beside it
- The old machinery assumes determinism, numeric inputs, a backtestable outcome and a release you control
- Forcing an LLM feature through unchanged yields validation metrics that do not measure anything relevant
- Declaring it out of scope because it does not fit is the failure mode supervisors ask about first
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.