Large language models may help identify drug safety signals in clinical notes, though their performance remains below thresholds required for clinical decision support.
Researchers evaluated three models — GPT-3.5, GPT-4 and GPT-4o — using clinical notes from 100 patients at Nashville, Tenn.-based Vanderbilt Health, 70 patients at the University of California—San Francisco and 272 patients from seven Roche-sponsored trials, according to an April 6 Vanderbilt news release.
For detecting immune-related adverse events at the patient level, GPT-4o achieved F1 scores of 56%, 66% and 62% across the respective datasets. The F1 score reflects how well a model balances correctly identifying real safety issues while avoiding false alarms. At the individual note level, the model reached an average F1 score of 57% across 667 notes.
An F1 score of 90% or more is considered excellent, while 80% or higher may support clinical decision-making.
Researchers said the models showed a tendency to overpredict adverse events but could help automate safety signal detection across sites and reduce reliance on manual chart review.
The study was published April 6 in eBioMedicine.