AI underperforms on simple number-crunching tasks that hospital administrators rely on to track admissions and apportion resources, Mount Sinai and Mayo Clinic researchers found.
The study’s authors analyzed nine large language models’ ability to perform two straightforward administrative tasks — tallying patients who meet a specific condition and narrowing records using multiple criteria — with data from 50,000 actual emergency department visits at New York City-based Mount Sinai Health System.
Basic prompts like “how many patients were admitted?” produced poor results across all models, according to the study published May 7 in PLOS Digital Health. Chain-of-thought reasoning, in which the AI is prompted to show its work, slightly improved accuracy, but performance dropped with larger data sets. A tool-based method, where models were asked to generate runnable code, significantly improved accuracy, though only for the strongest models.
“Our findings indicate that without using a tool-based strategy, current LLMs are unsuitable for standalone use even on minimally complex administrative tasks in clinical settings,” said Benjamin Glicksberg, PhD, associate professor of AI and human health at Icahn School of Medicine at Mount Sinai, in a May 7 news release. “Structured data tasks in clinical workflows will require agentic approaches that combine LLMs with code execution to ensure accuracy and consistency.”
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.