The AI Learning Hub Journal

Patient Data, De-identification, and Consent

What removing identifiers does, and what it does notde-identification lowers the chance of re-identification; it does not remove itWHAT COMES OUT — DIRECT IDENTIFIERSNames, contact details and addressesRecord, account and device identifiersExact dates tied to the individualPhotographs and free text that name someoneRemoving these is necessary, and it is the straightforward part.WHAT STAYS — QUASI-IDENTIFIERSAge band, sex, and rough locationAdmission and discharge timingA rare diagnosis or unusual care pathwayOccupation, employer, or referring serviceNone names anyone alone — in combination they can still single one out.RE-IDENTIFICATION RISK IS A SPECTRUM, NOT A SOLVED STATEharder to re-identifyeasier to re-identifyAggregate countsreleased aloneCoarsened bands,small cells suppressedFull record withquasi-identifiers intactRare profile that can bematched to outside dataNothing here is a fixed property — a new outside dataset moves any point rightwardsTHREE WORDS THAT ARE NOT SYNONYMSDE-IDENTIFIEDDirect identifiers removed. Riskis reduced, not removed, and theresidue is rarely measured.PSEUDONYMISEDA key still exists somewhere, sosomebody can reverse it. Stillpersonal data.ANONYMOUSRe-identification not reasonablypossible by anyone — a high bar,and claimed far too readily.AND IDENTIFIERS ARE NOT THE WHOLE OF THE PROTECTIONConsentWhat the patient agreed to, andwhether this use was within whatthey were actually told.TransparencyWhether patients can find out thattheir data is used this way atall, and by whom.Secondary-use limitsData gathered for care and reusedfor something else needs its ownbasis, not a lighter identifier.De-identified is a direction of travel, not a destination that has been reachedAsk what the data could be linked against, not only what was taken out of it
Educational orientation only — what is permitted, and on what basis, is set locally and differs by place

Two Regimes, Different Shapes

In the US, HIPAA governs protected health information held by covered entities and their business associates, structuring permitted uses and disclosures and imposing safeguards; notably, it attaches to categories of entity, so identical data can be in or out of scope depending on who holds it. In the EU, GDPR is a general data protection regime under which health data is special-category data requiring both a lawful basis and a specific condition for processing, with rights attaching to the individual regardless of who holds the data. The practical difference matters for AI: a US wellness app may hold health information outside HIPAA, while in the EU the same data carries special-category protection wherever it sits.

  • HIPAA attaches to covered entities and business associates — the holder determines scope
  • GDPR treats health data as special-category, requiring a lawful basis plus a processing condition
  • Identical data can be regulated differently in the US depending on who holds it
  • Cross-border processing and vendor hosting raise transfer questions on top of both

De-identification Has Limits

De-identification is the standard justification for using patient data in AI development, and it is weaker than it is usually presented. Removing direct identifiers does not remove the combinations that make records distinctive: rare conditions, unusual trajectories, precise dates, small geographies, and long longitudinal histories are all potentially identifying in combination with outside information. Imaging carries the additional problem that some scans contain reconstructable facial anatomy, and genomic data is intrinsically identifying and cannot be de-identified in any meaningful sense. Re-identification research has repeatedly demonstrated that supposedly anonymous health data can be linked back to individuals. Treat de-identified data as lower risk, not as outside the risk conversation.

  • Rare conditions, precise dates, small geographies, and long histories are identifying in combination
  • Some imaging permits facial reconstruction; genomic data is intrinsically identifying
  • Re-identification of supposedly anonymous health data has been demonstrated repeatedly
  • Treat de-identified as lower risk, not as non-personal

Consent, Transparency, and Secondary Use

Two distinct questions get conflated. First: does the patient know AI is involved in their care, and does that matter to them? Reasonable positions differ on whether every use requires disclosure, but the emerging expectation is that patients should be able to find out, and that anything materially affecting their care should be disclosed. Second: was their data used to build or improve the system? Secondary use for model development frequently rests on a basis the patient never actively considered, and vendor contracts sometimes permit training on institutional data in terms that clinical staff have never seen. Both questions deserve explicit institutional answers rather than being left to the procurement paperwork.

  • Disclosure of AI involvement in care and secondary use of data for training are separate questions
  • The emerging expectation is that patients can find out, with disclosure where care is materially affected
  • Check whether vendor terms permit training on your institution's data — clinical staff rarely see these clauses
  • Special-category data in the EU requires the processing condition to be identified explicitly, including for development

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.