OpenEvidence beats Claude, Gemini, GPT-5.5 in new physician-led study

Advertisement

OpenEvidence outperformed Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 on real-world clinical questions in a new physician-graded study, contradicting a June Nature Medicine paper that found general-purpose AI models beat specialized clinical tools.

The study, posted June 27 as a preprint on arXiv, has not yet been peer-reviewed. It had 149 practicing physicians across 36 states rate AI answers to 620 real point-of-care questions drawn from OpenEvidence’s platform, plus 187 questions from the HealthBench benchmark, judging responses on accuracy, clinical utility, source quality, verifiability and completeness.

OpenEvidence posted positive win margins on all five dimensions, with differences ranging from 25 to 39 percentage points over the three general-purpose models. Claude Opus 4.8 and Gemini 3.1 Pro scored near parity with each other, while GPT-5.5 recorded the lowest win rates on every axis.

The results stood in contrast to a Nature Medicine study published June 12, which found GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed OpenEvidence and Wolters Kluwer’s UpToDate Expert AI. OpenEvidence has since asked Nature Medicine to retract that study, alleging flawed methods; the journal has pointed the company to its formal rebuttal process instead.

The new paper’s authors attribute the divergence partly to design differences: Their evaluation used 149 physicians in specialty-matched, head-to-head comparisons, versus 12 clinicians at one institution using isolated rubric scoring in the earlier study. OpenEvidence codesigned the data collection plan, administered the survey and paid participating physicians, though the paper’s authors — based at University of California San Francisco, Boston-based Harvard Medical School and Stanford (Calif.) University, among others — report no affiliation with the company.

The competing results leave hospital and health system leaders with conflicting independent evidence on how general-purpose AI models compare with specialized clinical decision support tools as adoption of both accelerates.

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Register to Attend Webinar

AI literacy & clinical expertise: Keeping human judgment sharp as AI scales

Wednesday, August 12
11:00 AM - 12:00 PM CDT

Presenters: Lisa Ivanjack, MD, MHCM, FACP, ProvidenceJohn-Paul Mead, MD, Centralus HealthWilliam Gustin, MD, Moab Regional HospitalAmanda Heidemann, MD, FAAFP, FAMIA, Wolters Kluwer HealthYaw Fellin, Wolters Kluwer Health

Advertisement

Next Up in Artificial Intelligence

Advertisement