The AI Learning Hub Journal

Drift and Monitoring After Go-Live

Why a model that passed validation quietly stops workingnothing broke — the setting it was measured in moved, and the model did not move with itPOPULATION CHANGEWho comes through the doorshifts — case mix, ages,referral routes, seasons.EQUIPMENT OR PROTOCOLA new scanner, new settings,a revised protocol, anotherassay or another supplier.UPSTREAM DATA CHANGECoding practice shifts, afield gets repurposed, afeed changes its format.CLINICIAN BEHAVIOURPeople learn what the modelflags, and change what theysend it in the first place.A model in live use — the same weights, quietly meeting a different world each monthTHE MONITORING LOOP THAT CATCHES IT BEFORE ANYBODY ELSE DOES1Define the signalsWhich outputs, whichsubgroups, and what countsas a warning — agreed first.2Measure on live dataSample real cases and getoutcomes back, on a statedcadence rather than rumour.3Compare to the barAgainst the threshold setbefore go-live, not againstwhatever now looks fine.4Act on the findingRecalibrate, revalidate,narrow the scope, or takeit out of use altogether.the loop runs for as long as the model is in use — drift has no end dateContinue unchangedperformance is still insidethe limits agreed up frontRevalidate or recalibratewith the change documentedand the bar re-agreedRestrict or withdrawwhen it cannot be broughtback inside those limitsWithout the loop, the first signal that something changed arrives as a complaint or a harmEducational orientation only — what to monitor, and how often, is a local clinical decision
Validation is a photograph of one moment — monitoring is the only thing that tells you it still holds

Why Performance Decays

Model performance is not a fixed property; it is a relationship between a model and a world that keeps moving. Patient populations shift with demographics and referral changes. Clinical practice changes — new guidelines, new tests, new treatments. Upstream systems change: a record system upgrade alters field semantics, a scanner is replaced, a coding convention is updated, a form is redesigned. Disease patterns themselves change. Any of these can degrade a model without a single line of its code changing, and often without anyone connecting the change to the tool. A model deployed and left alone is not stable; it is unmonitored, and the difference only becomes visible when something goes wrong.

  • Populations, practice, upstream systems, and disease patterns all move underneath a static model
  • A record system upgrade or scanner replacement can degrade a model with no code change
  • Deployed and unmonitored is not the same as stable

What to Monitor

Monitoring works in layers. Input monitoring watches the data arriving — distributions, missingness rates, new codes, format changes — and is the earliest warning available, because it fires before outputs are affected. Output monitoring watches the model's own behaviour: flag rates, score distributions, and how often clinicians override it. Outcome monitoring, the hardest and most valuable, compares predictions against what actually happened, and requires an outcome ascertainment process that may not exist yet. All three should be broken down by subgroup, since aggregate stability can conceal degradation concentrated in one group. Rising override rates in particular are an informative early signal that clinicians have noticed something the metrics have not.

  • Input monitoring fires earliest — distributions, missingness, new codes, format changes
  • Output monitoring tracks flag rates, score distributions, and clinician override behaviour
  • Outcome monitoring is hardest and most valuable, and needs an ascertainment process
  • Disaggregate all three — aggregate stability hides subgroup-specific decay

The Feedback Loop Problem

Monitoring a deployed clinical model has a structural complication: once the model changes behaviour, it corrupts the data used to evaluate it. If a risk score prompts earlier intervention and the intervention works, the predicted outcomes stop occurring and the model looks less accurate — the intervention succeeded and the metric got worse. Conversely, if flagged patients receive more attention and therefore more testing, more findings will be confirmed in flagged patients regardless of model quality, making it look better than it is. Both effects are real and neither is easy to disentangle. It means monitoring must be interpreted with knowledge of what actions the output triggers, and that periodic independent evaluation remains necessary despite continuous monitoring.

  • Successful intervention removes the outcomes the model predicted, degrading apparent accuracy
  • Extra attention to flagged patients confirms more findings among them, inflating apparent accuracy
  • Monitoring must be read alongside the actions the output triggers
  • Continuous monitoring does not remove the need for periodic independent evaluation

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.