Foundation AI models trained on EHRs may inadvertently retain and expose sensitive patient information, according to a study by researchers at the Massachusetts Institute of Technology in Cambridge.
The study examined how clinical AI models can “memorize” individual patient data rather than generalize from broader trends. Researchers developed structured tests to determine how easily an attacker with partial knowledge — such as lab results or demographic details — could extract identifiable information from a model.
The team found that some patients, particularly those with rare conditions, may be more susceptible to privacy risks, even in de-identified datasets. While some disclosures — such as a patient’s age or gender — were seen as lower risk, others, including diagnoses related to HIV or substance use, were flagged as potentially harmful.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.