The AI Learning Hub Journal

What Imaging AI Genuinely Does Well

What imaging AI genuinely does well, and what it does notthe split is not about difficulty — it is about whether the question was specified in advanceGENUINELY STRONG ATDetection of specified findingsLooking for exactly the thing it was trained to lookfor, on every study, without getting tired.Measurement and quantificationSizes, volumes and densities, computed the same wayeach time rather than estimated by eye.Prioritising a worklistReordering the queue so studies that look urgentsurface earlier than they otherwise would.Flagging comparisonsPointing at the prior study and at where somethingappears to have changed since it was taken.NOT WHAT IT IS FORAnything outside its trained scopeA finding it was never taught to look for is, as faras the output is concerned, simply not present.Silence is not the same as a negative.Clinical contextThe history, the examination, the reason the studywas requested and what has already been tried allsit outside the pixels it was given.Judgement about the whole patientWhat this finding means for this person, what to donext, and whether it matters at all is a questionthe model was never asked to answer.Strong wherever the question was specified in advance, weak wherever it was notNothing in the output distinguishes “there is nothing there” from “nothing I was taught to see”WHAT THAT MEANS FOR THE PERSON READING THE STUDYA clean output is not a clean studyNothing flagged only means nothingon the list it was trained to flag.The reader still owns the restEverything outside the trained scopeis exactly as unassisted as before.Scope has to be written downIf nobody can state what it looksfor, nobody can state what it missed.Educational orientation only — what any given tool is cleared to do is stated by its own labelling
It answers the question it was trained on and stays silent on every other one — silence reads as reassurance

Four Tasks, Not One Capability

Imaging is the most mature clinical zone for AI, but the maturity is task-specific. Detection: flagging the possible presence of a defined finding. Quantification: measuring something a human would measure more slowly and less reproducibly — volumes, dimensions, densities, counts. Prioritisation: reordering a worklist so studies likelier to contain a time-critical finding are read sooner. Workflow support: automating the tedious parts of study handling, such as protocol selection, image quality checks, or pre-population of structured reports. These have different evidence bases and different consequences when they fail. "AI reads scans" is not a claim anyone should evaluate; the task is always narrower.

  • Detection, quantification, prioritisation, and workflow support are separate tasks with separate evidence
  • Quantification is often the most under-appreciated win: reproducible measurement, not judgement
  • Prioritisation changes the order of reading rather than the reading itself
  • Always ask which specific task, on which modality, for which finding

Why Narrow Tasks Are Where It Works

Imaging AI performs best where the task is well defined, the input is standardised, the finding has a consistent visual signature, and large labelled datasets exist. That description fits a specific set of problems well and fits general interpretation badly. A radiologist reading a study integrates clinical history, prior imaging, the referral question, incidental findings across every organ in the field, and an assessment of image quality. A model trained to detect one finding does one of those things. This is why single-finding tools can be genuinely useful and simultaneously nowhere near replacing interpretation — they are not partial versions of the same job, they are a different, much smaller job.

  • Best case: well-defined finding, standardised acquisition, consistent visual signature, large labelled data
  • A single-finding model does not do a scaled-down version of interpretation — it does a different task
  • Incidental findings outside the target finding remain entirely the reader's responsibility

What Follows for Evaluation

Because the task is narrow, the evaluation must be too. Ask which finding, on which modality, with which acquisition parameters, in which patient population, and against which reference standard the model was validated. Ask what the comparator was — an unaided reader, a reader with the tool, or nothing at all. Ask how the reference labels were produced, because a model validated against a single reader's labels inherits that reader's errors and cannot exceed them. And ask what happens to cases outside the intended finding, since a tool that is silent about everything else can still shape where attention goes. These questions are answerable by non-specialists and they filter out a great deal.

  • Which finding, which modality, which acquisition, which population, which reference standard
  • Labels derived from a single reader cap the model at that reader's accuracy, including their mistakes
  • The comparator matters more than the headline number — against nothing is not a comparison

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.