Researchers at Somerville, Mass.-based Mass General Brigham have developed BRIDGE, a multilingual benchmark that found large language models perform far better on medical licensing exams than actual patient-care tasks.
The benchmark evaluates how well LLMs understand clinical text — including the language used in EHRs, clinical case reports, and patient-doctor consultations — across nine languages. While the top-performing model scored as high as 92 on standardized medical exams, it earned just 44.8% on BRIDGE, exposing gaps in its grasp of nuanced clinical language, according to the findings published June 17 in Nature Biomedical Engineering.
“Unlike many existing medical AI benchmarks, BRIDGE focuses on real-world clinical data sources that better reflect the complexity of real-world care,” said senior author Jie Yang, PhD, of the Mass General Brigham Department of Medicine, in a June 17 news release. Dr. Yang added that the tool can help clinicians select the right AI while guiding developers toward better model performance.
Researchers used BRIDGE to evaluate 95 AI models across 14 clinical specialties, finding AI performance varied by specialty and language — a gap the tool could help close for non-English-speaking patients.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.