The AI Learning Hub Journal

How to Read an AI Study Critically

How strong is the evidence behind the claim?STRENGTH OF THE CLAIMInternal retrospectiveOld data, from the same placethe model was developedThe easiest test to passExternal retrospectiveOld data, from services themodel has never seenShows it travelsProspectiveRuns live in the real workflow,with real interruptionsNow the mess is includedRandomised / outcomeDoes care, or the patient,actually end up better?Strongest and rarestweakeststrongestDISCRIMINATION IS NOT UTILITYA model can separate cases well and still change nothing about what happens to the patient
Read any AI claim by asking which rung it stands on — most sit on the bottom one and are described as though they sat on the top

Retrospective vs Prospective

Most published healthcare AI evidence is retrospective: the model is run over data already collected, with outcomes already known. This is cheap, fast, and systematically optimistic. Retrospective datasets are cleaned, complete, and often assembled with the study question in mind; cases with missing data or ambiguous labels tend to be excluded, removing exactly the cases that are hard in practice. Prospective evaluation runs the model on patients as they present, in real time, on real data pipelines, with real missingness. Performance almost always drops. The drop is not a scandal — it is the difference between a controlled measurement and reality. A tool with only retrospective evidence has not yet been tested in the condition it will be used in.

  • Retrospective data is cleaned and curated in ways that remove the hardest cases
  • Prospective evaluation exposes real missingness, timing, and pipeline behaviour
  • Expect performance to drop prospectively — the question is by how much, and whether anyone measured
  • Retrospective-only evidence means the tool has not been tested in its operating condition

Internal vs External Validation

Internal validation holds out part of the same dataset for testing. It guards against one specific error — memorising the training examples — and nothing else. Because the held-out data shares the same sites, equipment, populations, coding conventions, and label definitions, it cannot detect any of the shifts that break models in deployment. External validation uses data from different institutions, ideally different regions and health systems, collected independently. It is the only design that tests generalisation. When you see impressive numbers, the first question is which kind of validation produced them. Internal-only results, however strong, tell you the model learned the dataset — a necessary condition and a very weak one.

  • Internal validation only rules out memorising the training set
  • Shared sites, equipment, coding, and label definitions mean shared blind spots
  • External validation across independent institutions is the only real test of generalisation
  • Strong internal-only numbers are a necessary but very weak result

Design Details That Quietly Inflate Results

Several recurring design choices flatter results without being dishonest. Case-control assembly — collecting clear positives and clear negatives — removes the ambiguous middle where clinical difficulty lives and inflates apparent discrimination. Data leakage, such as multiple studies from the same patient split across training and test sets, lets the model recognise the patient rather than the pathology. Labels derived from the very clinicians the model is compared against build in circularity. Unclear handling of missing data can encode the fact that a test was ordered, which is itself informative. Reporting guidelines specifically for AI in health now exist, as extensions of established clinical study and prediction-model standards; checking whether a study follows one is a fast quality filter.

  • Case-control assembly removes the ambiguous cases that make the task hard
  • Patient-level leakage across splits lets the model recognise people, not disease
  • Labels drawn from the comparator clinicians build circularity into the result
  • AI-specific reporting guidelines exist — whether a study follows one is a quick quality signal
  • Named examples: CONSORT-AI and SPIRIT-AI for trials and protocols, and the TRIPOD standard's AI extension for prediction models

Try It Yourself

These checks are worth more applied once than read twice. Take a real claim someone is making in your field and put it through them.

◆ Try it yourself

Take one publicly available study or vendor document making a performance claim about an AI tool in your field — published papers, preprints, or public white papers only, never internal or patient data. Work through it and answer in writing: retrospective or prospective, internal or external validation, what the comparator was, how the reference labels were produced, and which cases were excluded.

Claim being made:
Retrospective or prospective:
Internal or external validation, and how many sites:
Comparator used:
How reference labels were produced:
Cases excluded from the dataset:
What this evidence does NOT establish:
How you'll know it worked
  • You answered all five questions from the document itself, or recorded that it does not say
  • You can state in one sentence what the study does not establish about your setting
  • You could defend that summary to the person promoting the tool

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.