The study, published Oct. 8 in Nature Communications, compiled a set of 1,000 of UCSF Health’s ED visits with the same ratio of “yes” to “no” responses for decision on admission, radiology and antibiotics. Researchers entered the physician’s notes on each patient’s symptoms and examination findings into ChatGPT-3.5 and ChatGPT-4.
The study found ChatGPT tended to recommend services more often than needed. ChatGPT-4 was also 8% less accurate than resident physicians, and ChatGPT-3.5 was 24% less accurate.
“This is a valuable message to clinicians not to blindly trust these models,” the study’s lead author, Chris Williams, MD, said in the release. “ChatGPT can answer medical exam questions and help draft clinical notes, but it’s not currently designed for situations that call for multiple considerations, like the situations in an emergency department.”
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.