Patient Data, De-identification, and Consent
Two Regimes, Different Shapes
In the US, HIPAA governs protected health information held by covered entities and their business associates, structuring permitted uses and disclosures and imposing safeguards; notably, it attaches to categories of entity, so identical data can be in or out of scope depending on who holds it. In the EU, GDPR is a general data protection regime under which health data is special-category data requiring both a lawful basis and a specific condition for processing, with rights attaching to the individual regardless of who holds the data. The practical difference matters for AI: a US wellness app may hold health information outside HIPAA, while in the EU the same data carries special-category protection wherever it sits.
- HIPAA attaches to covered entities and business associates — the holder determines scope
- GDPR treats health data as special-category, requiring a lawful basis plus a processing condition
- Identical data can be regulated differently in the US depending on who holds it
- Cross-border processing and vendor hosting raise transfer questions on top of both
De-identification Has Limits
De-identification is the standard justification for using patient data in AI development, and it is weaker than it is usually presented. Removing direct identifiers does not remove the combinations that make records distinctive: rare conditions, unusual trajectories, precise dates, small geographies, and long longitudinal histories are all potentially identifying in combination with outside information. Imaging carries the additional problem that some scans contain reconstructable facial anatomy, and genomic data is intrinsically identifying and cannot be de-identified in any meaningful sense. Re-identification research has repeatedly demonstrated that supposedly anonymous health data can be linked back to individuals. Treat de-identified data as lower risk, not as outside the risk conversation.
- Rare conditions, precise dates, small geographies, and long histories are identifying in combination
- Some imaging permits facial reconstruction; genomic data is intrinsically identifying
- Re-identification of supposedly anonymous health data has been demonstrated repeatedly
- Treat de-identified as lower risk, not as non-personal
Consent, Transparency, and Secondary Use
Two distinct questions get conflated. First: does the patient know AI is involved in their care, and does that matter to them? Reasonable positions differ on whether every use requires disclosure, but the emerging expectation is that patients should be able to find out, and that anything materially affecting their care should be disclosed. Second: was their data used to build or improve the system? Secondary use for model development frequently rests on a basis the patient never actively considered, and vendor contracts sometimes permit training on institutional data in terms that clinical staff have never seen. Both questions deserve explicit institutional answers rather than being left to the procurement paperwork.
- Disclosure of AI involvement in care and secondary use of data for training are separate questions
- The emerging expectation is that patients can find out, with disclosure where care is materially affected
- Check whether vendor terms permit training on your institution's data — clinical staff rarely see these clauses
- Special-category data in the EU requires the processing condition to be identified explicitly, including for development
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.