Dataset Bias and Documented Harms
Health Data Records Access, Not Just Biology
Clinical datasets are records of who reached care, what was tested, what was coded, and how it was documented. All of those are shaped by access, insurance, geography, language, clinician behaviour, and historical inequity. A model trained on that data learns those patterns as faithfully as it learns physiology, because it cannot distinguish between them. This is why bias in health AI is not primarily a question of intent or of a badly chosen algorithm. It is a structural property of learning from records generated by a system that already treats people differently. Assuming a dataset is neutral because it is clinical is the mistake that produces the documented harms.
- Clinical data encodes access, testing patterns, coding practice, and documentation behaviour
- A model cannot separate a pattern of disease from a pattern of unequal care
- Bias here is structural, not a matter of intent or algorithm choice
Proxy Variables: The Cost-as-Need Case
The most instructive documented case in health AI concerns a widely used US care-management algorithm that predicted future healthcare costs as a stand-in for future healthcare need, then used that prediction to select patients for extra support. The logic seems reasonable: sicker patients cost more. But less money is historically spent on Black patients at equivalent levels of illness, for reasons of access and structural inequity. The model learned the spending pattern accurately, and in doing so systematically underestimated need among Black patients, directing support away from people who were equally or more unwell. Nothing about the model was technically wrong. The proxy was wrong, and the harm ran at population scale before anyone examined it.
- Cost was used as a proxy for need; historical spending is not equal across groups at equal illness
- The model learned the spending pattern correctly and produced inequitable allocation as a result
- Interrogate every proxy: what was actually measured, and what was it standing in for?
- This harm was invisible in standard performance metrics — it required a fairness-specific audit
- The case is documented in a peer-reviewed 2019 study in Science by Obermeyer and colleagues
Measurement, Devices, and Labels
Bias also enters through the instruments. Optical measurement devices have documented accuracy differences by skin pigmentation, meaning the inputs a model learns from are themselves systematically less accurate for some patients; a model trained on those readings inherits the discrepancy and can compound it. Dermatology image datasets have long been noted as skewed towards lighter skin, limiting what models learn about presentations on darker skin. Label bias is subtler still: if the label is "was diagnosed" rather than "had the condition", the model learns historical diagnostic patterns including under-diagnosis in specific groups, and reproduces them confidently. In each case the model is faithfully reproducing a flaw upstream of itself.
- Device measurement accuracy can vary by skin pigmentation — biased inputs produce biased models
- Skewed image datasets limit what a model learns about presentations in underrepresented groups
- Labels record what was diagnosed, not what was present — under-diagnosis is learned and reproduced
- The model is usually faithful to the data; the defect sits upstream
What to Actually Do About It
Fairness cannot be established from an aggregate metric, so the minimum practical step is to require performance disaggregated by the groups that matter in your population, and to treat the absence of that breakdown as a finding rather than a gap in the paperwork. Beyond that: interrogate what the outcome variable really represents and whether it is a proxy; check whether the development population resembles yours; ask how missing data was handled, since missingness itself is often patterned by group; and ensure the deployment includes a mechanism to detect differential impact after go-live, because the harms above were all discovered late by people who went looking.
- Require subgroup performance; treat its absence as evidence, not an administrative omission
- Interrogate the outcome variable — is it the thing you care about, or a proxy for it?
- Missingness is patterned by group and can carry the bias by itself
- Build differential-impact monitoring into deployment; these harms are found by looking
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.