What Imaging AI Genuinely Does Well
Four Tasks, Not One Capability
Imaging is the most mature clinical zone for AI, but the maturity is task-specific. Detection: flagging the possible presence of a defined finding. Quantification: measuring something a human would measure more slowly and less reproducibly — volumes, dimensions, densities, counts. Prioritisation: reordering a worklist so studies likelier to contain a time-critical finding are read sooner. Workflow support: automating the tedious parts of study handling, such as protocol selection, image quality checks, or pre-population of structured reports. These have different evidence bases and different consequences when they fail. "AI reads scans" is not a claim anyone should evaluate; the task is always narrower.
- Detection, quantification, prioritisation, and workflow support are separate tasks with separate evidence
- Quantification is often the most under-appreciated win: reproducible measurement, not judgement
- Prioritisation changes the order of reading rather than the reading itself
- Always ask which specific task, on which modality, for which finding
Why Narrow Tasks Are Where It Works
Imaging AI performs best where the task is well defined, the input is standardised, the finding has a consistent visual signature, and large labelled datasets exist. That description fits a specific set of problems well and fits general interpretation badly. A radiologist reading a study integrates clinical history, prior imaging, the referral question, incidental findings across every organ in the field, and an assessment of image quality. A model trained to detect one finding does one of those things. This is why single-finding tools can be genuinely useful and simultaneously nowhere near replacing interpretation — they are not partial versions of the same job, they are a different, much smaller job.
- Best case: well-defined finding, standardised acquisition, consistent visual signature, large labelled data
- A single-finding model does not do a scaled-down version of interpretation — it does a different task
- Incidental findings outside the target finding remain entirely the reader's responsibility
What Follows for Evaluation
Because the task is narrow, the evaluation must be too. Ask which finding, on which modality, with which acquisition parameters, in which patient population, and against which reference standard the model was validated. Ask what the comparator was — an unaided reader, a reader with the tool, or nothing at all. Ask how the reference labels were produced, because a model validated against a single reader's labels inherits that reader's errors and cannot exceed them. And ask what happens to cases outside the intended finding, since a tool that is silent about everything else can still shape where attention goes. These questions are answerable by non-specialists and they filter out a great deal.
- Which finding, which modality, which acquisition, which population, which reference standard
- Labels derived from a single reader cap the model at that reader's accuracy, including their mistakes
- The comparator matters more than the headline number — against nothing is not a comparison
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.