AI could streamline drug safety detection in EHRs: Study

Advertisement

Large language models may help identify drug safety signals in clinical notes, though their performance remains below thresholds required for clinical decision support.

Researchers evaluated three models — GPT-3.5, GPT-4 and GPT-4o — using clinical notes from 100 patients at Nashville, Tenn.-based Vanderbilt Health, 70 patients at the University of California—San Francisco and 272 patients from seven Roche-sponsored trials, according to an April 6 Vanderbilt news release.

For detecting immune-related adverse events at the patient level, GPT-4o achieved F1 scores of 56%, 66% and 62% across the respective datasets. The F1 score reflects how well a model balances correctly identifying real safety issues while avoiding false alarms. At the individual note level, the model reached an average F1 score of 57% across 667 notes.

An F1 score of 90% or more is considered excellent, while 80% or higher may support clinical decision-making.

Researchers said the models showed a tendency to overpredict adverse events but could help automate safety signal detection across sites and reduce reliance on manual chart review.

The study was published April 6 in eBioMedicine.

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Register to Attend Webinar

The Performance Gap Your Pharmacy Training Metrics Can’t See

Monday, August 10
12:00 PM - 1:00 PM CDT

Presenters: Michael Alexander, AudirieDr. Tina Moen, PharmD, Colibri Healthcare

Advertisement

Next Up in Pharmacy

Advertisement