The AI Learning Hub Journal

Dataset Bias and Documented Harms

How bias gets into health data, and what it does once therefour ordinary, unglamorous routes in — and one reporting habit that makes them visibleUNDER-REPRESENTATIONWhole groups appear rarelyor not at all, so nothingever forced the model toget them right.PROXY VARIABLESSomething easy to recordstands in for the thingactually being predicted,and carries its history.MEASUREMENT DIFFERENCESDevices, settings and skintones change the signalitself, long before anymodel has seen it.LABEL NOISEDocumentation practicevaries between people andsites, so the same case iswritten down differently.None of these announce themselves — they arrive as a model that works less well for some peopleWHAT THE HEADLINE NUMBER SHOWSOne overall figure, averaged across everybody whohappened to be in the evaluation set. It looks fine,and it is not wrong — it is simply not an answer.OVERALL — LOOKS ACCEPTABLEThe average is held up by whoever the data had most of.THE SAME MODEL, SPLIT BY SUBGROUPLargest group in the training dataAnother well-represented groupDifferent device or settingUnder-represented groupRarely and unevenly documentedHOW YOU FIND IT BEFORE SOMEBODY ELSE DOESPre-specify the subgroupsDecide before you measure, so thecut is not chosen after the fact.Report every subgroup, alwaysA single headline figure is the onenumber able to hide all of this.Set the bar on the worst oneIf the weakest subgroup fails thenthe model fails, not the average.Unequal performance is the symptom; the data collection is where it was decidedA model can only be as fair as the record it learned from, and records were never neutralEducational orientation only — which subgroups matter is a question for the setting and its people
An average is a summary of everyone and a description of no one — bias hides in exactly that gap

Health Data Records Access, Not Just Biology

Clinical datasets are records of who reached care, what was tested, what was coded, and how it was documented. All of those are shaped by access, insurance, geography, language, clinician behaviour, and historical inequity. A model trained on that data learns those patterns as faithfully as it learns physiology, because it cannot distinguish between them. This is why bias in health AI is not primarily a question of intent or of a badly chosen algorithm. It is a structural property of learning from records generated by a system that already treats people differently. Assuming a dataset is neutral because it is clinical is the mistake that produces the documented harms.

  • Clinical data encodes access, testing patterns, coding practice, and documentation behaviour
  • A model cannot separate a pattern of disease from a pattern of unequal care
  • Bias here is structural, not a matter of intent or algorithm choice

Proxy Variables: The Cost-as-Need Case

The most instructive documented case in health AI concerns a widely used US care-management algorithm that predicted future healthcare costs as a stand-in for future healthcare need, then used that prediction to select patients for extra support. The logic seems reasonable: sicker patients cost more. But less money is historically spent on Black patients at equivalent levels of illness, for reasons of access and structural inequity. The model learned the spending pattern accurately, and in doing so systematically underestimated need among Black patients, directing support away from people who were equally or more unwell. Nothing about the model was technically wrong. The proxy was wrong, and the harm ran at population scale before anyone examined it.

  • Cost was used as a proxy for need; historical spending is not equal across groups at equal illness
  • The model learned the spending pattern correctly and produced inequitable allocation as a result
  • Interrogate every proxy: what was actually measured, and what was it standing in for?
  • This harm was invisible in standard performance metrics — it required a fairness-specific audit
  • The case is documented in a peer-reviewed 2019 study in Science by Obermeyer and colleagues

Measurement, Devices, and Labels

Bias also enters through the instruments. Optical measurement devices have documented accuracy differences by skin pigmentation, meaning the inputs a model learns from are themselves systematically less accurate for some patients; a model trained on those readings inherits the discrepancy and can compound it. Dermatology image datasets have long been noted as skewed towards lighter skin, limiting what models learn about presentations on darker skin. Label bias is subtler still: if the label is "was diagnosed" rather than "had the condition", the model learns historical diagnostic patterns including under-diagnosis in specific groups, and reproduces them confidently. In each case the model is faithfully reproducing a flaw upstream of itself.

  • Device measurement accuracy can vary by skin pigmentation — biased inputs produce biased models
  • Skewed image datasets limit what a model learns about presentations in underrepresented groups
  • Labels record what was diagnosed, not what was present — under-diagnosis is learned and reproduced
  • The model is usually faithful to the data; the defect sits upstream

What to Actually Do About It

Fairness cannot be established from an aggregate metric, so the minimum practical step is to require performance disaggregated by the groups that matter in your population, and to treat the absence of that breakdown as a finding rather than a gap in the paperwork. Beyond that: interrogate what the outcome variable really represents and whether it is a proxy; check whether the development population resembles yours; ask how missing data was handled, since missingness itself is often patterned by group; and ensure the deployment includes a mechanism to detect differential impact after go-live, because the harms above were all discovered late by people who went looking.

  • Require subgroup performance; treat its absence as evidence, not an administrative omission
  • Interrogate the outcome variable — is it the thing you care about, or a proxy for it?
  • Missingness is patterned by group and can carry the bias by itself
  • Build differential-impact monitoring into deployment; these harms are found by looking

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.