Drift and Monitoring After Go-Live
Why Performance Decays
Model performance is not a fixed property; it is a relationship between a model and a world that keeps moving. Patient populations shift with demographics and referral changes. Clinical practice changes — new guidelines, new tests, new treatments. Upstream systems change: a record system upgrade alters field semantics, a scanner is replaced, a coding convention is updated, a form is redesigned. Disease patterns themselves change. Any of these can degrade a model without a single line of its code changing, and often without anyone connecting the change to the tool. A model deployed and left alone is not stable; it is unmonitored, and the difference only becomes visible when something goes wrong.
- Populations, practice, upstream systems, and disease patterns all move underneath a static model
- A record system upgrade or scanner replacement can degrade a model with no code change
- Deployed and unmonitored is not the same as stable
What to Monitor
Monitoring works in layers. Input monitoring watches the data arriving — distributions, missingness rates, new codes, format changes — and is the earliest warning available, because it fires before outputs are affected. Output monitoring watches the model's own behaviour: flag rates, score distributions, and how often clinicians override it. Outcome monitoring, the hardest and most valuable, compares predictions against what actually happened, and requires an outcome ascertainment process that may not exist yet. All three should be broken down by subgroup, since aggregate stability can conceal degradation concentrated in one group. Rising override rates in particular are an informative early signal that clinicians have noticed something the metrics have not.
- Input monitoring fires earliest — distributions, missingness, new codes, format changes
- Output monitoring tracks flag rates, score distributions, and clinician override behaviour
- Outcome monitoring is hardest and most valuable, and needs an ascertainment process
- Disaggregate all three — aggregate stability hides subgroup-specific decay
The Feedback Loop Problem
Monitoring a deployed clinical model has a structural complication: once the model changes behaviour, it corrupts the data used to evaluate it. If a risk score prompts earlier intervention and the intervention works, the predicted outcomes stop occurring and the model looks less accurate — the intervention succeeded and the metric got worse. Conversely, if flagged patients receive more attention and therefore more testing, more findings will be confirmed in flagged patients regardless of model quality, making it look better than it is. Both effects are real and neither is easy to disentangle. It means monitoring must be interpreted with knowledge of what actions the output triggers, and that periodic independent evaluation remains necessary despite continuous monitoring.
- Successful intervention removes the outcomes the model predicted, degrading apparent accuracy
- Extra attention to flagged patients confirms more findings among them, inflating apparent accuracy
- Monitoring must be read alongside the actions the output triggers
- Continuous monitoring does not remove the need for periodic independent evaluation
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.