AI has outperformed physicians on emergency diagnoses, flagged pancreatic cancer up to three years early and beat specialized clinical software on medical benchmarks across several studies Becker’s has covered in recent months.
But the studies raise as many questions as they answer for health system leaders. Most of the strongest results came from retrospective, simulated or single-site studies, leaving uncertainty about how these tools perform in routine clinical practice — a validation gap some authors themselves acknowledged.
At the same time, AI is still largely framed as a tool for generating leads that clinicians review, leaving accountability unclear when models miss or misidentify findings. And with some performance claims coming from vendors rather than independent peer-reviewed research, leaders must carefully distinguish between validated evidence and company-reported results before making investment decisions.
Here are eight studies about AI in diagnosis and clinical reasoning that Becker’s has reported on in the past two months:
1. Researchers from Boston Children’s Hospital, Cambridge, Mass.-based Harvard University and OpenAI confirmed 18 diagnoses after using an AI-assisted workflow to reanalyze 376 previously unsolved rare disease cases, in a study published June 18 in NEJM AI.
2. OpenAI said June 18 that a physician panel rated its new GPT-5.5 Instant model higher than answers written by physicians — who had unlimited time and internet access — across 3,500 reviewed responses, scoring it higher on accuracy, communication, completeness and health decision helpfulness.
3. General-purpose large language models from OpenAI, Google and Anthropic outperformed specialized clinical AI tools across every medical benchmark in a study published June 12 in Nature Medicine, though the clinical AI developers disputed the findings.
4. Researchers at Cleveland Clinic and Pittsburgh-based Carnegie Mellon University developed an AI system that interprets cardiac MRI scans without manually labeled training data, according to a May 21 announcement and a study in Nature Communications.
5. An OpenAI o1 model outperformed two human physicians on emergency department diagnoses, identifying the exact or a very close diagnosis more often across 76 cases at Beth Israel Deaconess Medical Center, according to an April 30 study in Science.
6. Researchers at Worcester-based UMass Chan Medical School tested a real-time AI tool to diagnose cholangiocarcinoma during live procedures in the first-in-human trial of its kind, according to a study published April 9 in Clinical Gastroenterology and Hepatology.
7. An AI model developed by researchers at Rochester, Minn.-based Mayo Clinic detected pancreatic cancer on abdominal CT scans taken up to three years before clinical diagnosis, according to a study published April 28 in Gut.
8. A conversational AI tool developed by researchers at University of California San Diego guided patients through self-triage using established medical protocols, selecting the correct medical flowchart about 84% of the time and following decision-making steps with more than 99% accuracy across more than 30,000 simulated patient interactions, according to a study published April 23 in Nature Health.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.