The AI Learning Hub Journal

Sensitivity, Specificity, and Population

The trade you are actually making when you set a thresholdshape only — no numbers, because the numbers depend entirely on the settingOPERATING THRESHOLDpeople withoutthe conditionpeople withthe conditiontest reads more negativetest reads more positiveMISSED CASESThe condition is present, the test does not flag it,and it may only be found later, if at all.FALSE ALARMSFlagged, then found not to have it after all —worry, cost, and a follow-up that was not needed.WHAT EACH MISTAKE COSTSA missed case that was treatable and serious pullsthe threshold one way. A false alarm that leads tosomething invasive and frightening pulls it back.HOW COMMON IT IS IN THIS GROUPThe same test behaves differently in a group wherethe condition is common than in one where it israre — which is why screening is not diagnosis.There is no setting that is simply better — only one chosen for a stated purpose and population
Moving the line does not make a test better — it decides which of the two mistakes you would rather make

One Dial, Two Costs

Every detection model has an internal score, and someone chooses where to cut it. Move the cut one way and the model catches more true cases while also flagging more that are not there. Move it the other way and false alarms fall while genuine cases are missed. There is no setting that improves both; there is only a choice about which error you prefer, and that choice belongs to the clinical context, not the model. The costs are asymmetric and specific: a missed finding of one kind is catastrophic, while a false alarm of another kind means an unnecessary procedure, anxiety, and cost. A single headline accuracy figure hides this choice entirely, which is exactly why it gets quoted.

  • Sensitivity and specificity trade against each other — no threshold improves both
  • The right balance is a clinical and contextual judgement, not a technical default
  • Headline accuracy conceals which error the vendor chose to minimise

Prevalence Changes What a Positive Means

Sensitivity and specificity are properties of the model. What a clinician cares about is different: given a positive flag, how likely is the finding actually present? That depends on how common the condition is in the population being scanned. Run the same model in a specialist referral setting where the condition is common, and most positives are real. Run it in general screening where the condition is rare, and the same model produces far more false positives than true ones, even with unchanged sensitivity and specificity. Nothing about the model has changed. This is the single most common misreading of imaging AI performance claims, and it fully explains why models look strong in development and disappointing in screening.

  • Sensitivity and specificity are model properties; predictive value depends on prevalence in your population
  • In low-prevalence settings, false positives can outnumber true positives with no change to the model
  • Always ask what the prevalence was in the validation set versus your intended setting

Tuned for Them, Not for You

Follow the logic through and the consequence is uncomfortable. A model with an operating point chosen for one setting is, in a real sense, a different tool in another. Age distribution, comorbidity mix, referral pathways, scanning indications, and local practice all shift prevalence and case difficulty. A tool developed on referral-centre data will encounter easier, rarer, and differently distributed cases in community practice. The threshold that made it useful there may make it a nuisance here, or unsafe. This is not a defect to be fixed by a better model; it is an intrinsic property of thresholded prediction, and it is why local validation appears in every serious deployment framework.

  • An operating point tuned for one population is a different tool in another
  • Case mix, referral pathway, and indication all shift prevalence and difficulty
  • This is intrinsic to thresholded prediction, not a bug a better model removes
  • It is the core technical reason local validation is mandatory rather than optional

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.