The AI Learning Hub Journal

Distribution Shift: Scanners, Sites, and People

The same model, a different settingWHERE IT WAS BUILTOne service, its scannersits protocols and staffits patient mixmeasured performance lives hereThe model travelscarrying patterns from where itlearned, intended or notWHERE IT GETS USEDAnother service, other kitother protocols and staffanother patient mixperformance is unknown until testedWHAT ACTUALLY CHANGED BETWEEN THE TWOEquipmentDifferent makes, ages andsettings; different image qualityProtocolsHow the test is ordered, run,labelled and reportedPopulationAge, comorbidity, body habitus,ancestry, referral routePrevalenceHow common the finding is herechanges what a flag meansBEFORE DEPLOYMENT — VALIDATE LOCALLYTest on your own cases, not the vendor setCompare against how the work is done todayAgree in advance what good enough here meansAFTER DEPLOYMENT — WATCH FOR DRIFTEquipment, staffing and case mix keep movingTrack inputs and outputs, not just uptimeRe-check whenever anything upstream changesEvidence from somewhere else is a reason to test here — never a substitute for testing here
A model learns its training environment as well as its task — performance is a property of the pairing, not of the model alone

Acquisition Shift

Medical images are not neutral records; they are the product of a specific device, protocol, and set of settings. Manufacturer, model, generation, field strength, dose, reconstruction algorithm, slice thickness, contrast timing, and positioning conventions all leave systematic signatures in the pixels. A model trained on one site's equipment has learned those signatures along with the pathology. Move it to a site with different scanners or protocols and performance can degrade without any visible change in image quality to a human reader. This is why external validation on different equipment is a distinct requirement from validation on more patients: more data from the same scanners does not test the thing that breaks.

  • Device, protocol, reconstruction, and positioning conventions leave systematic signatures in the data
  • Degradation can occur with no perceptible image quality change to a human reader
  • More data from the same equipment does not substitute for validation on different equipment

Population Shift

The second shift is in who is being imaged. Training populations differ from deployment populations in age, sex, ethnicity, body habitus, disease severity, comorbidity burden, and the reason the study was ordered. Each of these can change how a finding appears or how often it occurs. Underrepresentation in training data means the model has had less opportunity to learn the appearance of the finding in that group, and performance is typically worse — but reported average performance will not reveal it unless results are broken down by subgroup. If a validation report gives you one number and no subgroup breakdown, you cannot tell whether the tool works equally for your patients, and you should say so.

  • Age, sex, ethnicity, habitus, severity, comorbidity, and indication all shift between development and deployment
  • Underrepresented groups typically see worse performance, hidden inside the average
  • No subgroup breakdown means no evidence of equitable performance — treat that as a finding

Shortcut Learning

The most instructive failure mode is when a model achieves strong performance by learning something correlated with the diagnosis rather than the diagnosis itself. Documented examples across medical imaging research include models keying on scanner or site markers, on burned-in text and laterality tokens, on the presence of equipment such as drains or lines that indicate a patient is already being treated, or on positioning differences between how sick and well patients are imaged. Each shortcut works beautifully in the dataset that contains it and fails the moment the correlation breaks. It is also why a model can be highly accurate and clinically worthless at the same time, and why performance dropping sharply at a new site is a signal to investigate rather than to retrain blindly.

  • Models can learn site markers, burned-in text, visible equipment, or positioning instead of pathology
  • Shortcuts produce excellent in-dataset performance and collapse when the correlation breaks
  • A sharp drop at a new site is a signal to investigate what the model was actually using
  • High accuracy and clinical worthlessness are entirely compatible

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.