The AI Learning Hub Journal

Regression Suites That Survive Model Upgrades

The regression suite in CIA CHANGE LANDSPrompt editModel upgradeRetrieval tweakEVAL SUITE RUNSEvery case, every graderN repeats per caseModel version + prompt hash pinnedLatency and cost captured tooRuns on the PR, not after mergeCOMPARE TO BASELINEScore delta per metricPer-case pass to fail flipsNew failures listed by nameWins reported as well as lossesA diff, not a single numberRELEASE GATEregressionsblock the mergeTHE BASELINELast release scores, per case and per metricStored with model version, prompt hash, dataset hashRe-baselined deliberately — never silentlySHIP ITNo regression past the bandBaseline moves to this runDeltas posted on the PRBLOCK AND TRIAGENamed cases regressedFix, or accept with a reasonNever re-run until greenNON-DETERMINISM — the same input does not give you the same outputRepeat each caserun N times, report meanand spread, not one sampleGate on a bandfail on a drop beyond theband, not on any wobblePin what you cantemperature, seed, modelversion, dataset hashWithout a stored baseline there is no regression — only a score nobody can argue with
CI for AI is a diff against a baseline, not a green tick — and a noisy system needs a tolerance band before it needs a stricter gate

Assert on Outcomes, Not Phrasing

A regression suite earns its keep by staying valid across model changes, and the single biggest threat to that is assertions coupled to surface form. Exact-match checks against a reference paragraph, expected orderings of list items, specific phrasings in an explanation — all of these fail on a new model that is behaving perfectly well, generating a flood of false regressions precisely when you need to read a diff carefully. The discipline is to assert the properties that define correctness and nothing more: the required facts appear, the schema validates, the computed figure matches, the file compiles, the correct tool was called with valid arguments, the forbidden action never occurred. Write graders as though the wording will change on every run, because eventually it will.

  • Assert required content, structure, and post-conditions — never exact prose
  • Surface-coupled assertions produce false regressions exactly when clarity matters most
  • Prefer post-conditions on the world over inspection of the text describing them

Separate the Contract from the Tuning

Prompts accumulate model-specific tuning — phrasings that worked around one family's quirks, formatting tricks, few-shot examples chosen for a particular tokenizer's habits. This is legitimate engineering that also becomes migration debt, because it is indistinguishable from essential behaviour once it is all one blob of text. Structure the prompt so the two are visibly separate: the contract — task definition, required output schema, hard constraints, domain rules — kept apart from the model-specific layer of phrasing and examples. Your suite should test the contract. When you evaluate a new model, you then re-tune the model-specific layer and hold the contract fixed, which turns a rewrite into a swap and makes it obvious whether a failure is a genuine capability gap or an artefact of tuning aimed at a different model.

  • Keep task contract and model-specific tuning in separate, composable pieces
  • Evals target the contract; the tuning layer is expected to change per model
  • A failure in the contract is a capability gap; a failure in tuning is a config task
  • Run the suite against at least two model families to keep hidden coupling visible

The Upgrade Drill

Treat model migration as a rehearsed procedure rather than an emergency. Run the full suite against the candidate model and read results by slice, not in aggregate. Expect a mix: genuine improvements, genuine regressions, and false regressions from cases that were always fragile — and investigate a sample of each rather than assuming. Where the new model behaves differently but arguably better, the case may be wrong; the correct response is to fix the case, and to note that a case whose expected answer changes because a better model appeared was probably over-specified. Then shadow the candidate on live traffic without serving its output, compare on production inputs, and roll out progressively with the eval suite running against real traffic. Do this on a schedule, not only when a deprecation notice forces it.

  • Read migration results by slice; aggregates hide the trade you are actually making
  • Investigate regressions and improvements alike — some regressions are bad cases
  • Shadow-run on live traffic before serving, then roll out progressively
  • Rehearse upgrades on a cadence so the forced one is routine

Keep the Suite from Rotting

Regression suites decay in predictable ways. They accumulate cases nobody understands, written for a bug fixed two rewrites ago, that now fail for reasons unrelated to their original purpose. They fill with near-duplicates from repeated incidents. They grow slow enough that people stop running them before merging. Counter all three deliberately: every case carries a one-line note on why it exists and what it protects, cases that have not failed in a long time get reviewed for whether they still discriminate, duplicates are merged, and the suite is split into a fast pre-merge tier and a comprehensive nightly tier. A regression suite is a code asset with maintenance costs, and the teams that treat it as write-once end up with a slow suite full of noise that everyone has learned to ignore.

  • Every case documents what it protects — undocumented cases become unremovable
  • Review long-passing cases for continued relevance; merge duplicates
  • Split fast pre-merge and comprehensive nightly tiers to protect iteration speed

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.