General-purpose large language models from OpenAI, Google and Anthropic outperformed specialized clinical AI tools across every medical benchmark in a study published in Nature Medicine.
Researchers tested OpenEvidence and Wolters Kluwer’s UpToDate Expert AI against GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 in three stages: 500 MedQA knowledge questions, 500 HealthBench items measuring clinician alignment, and 100 real clinical queries drawn from physician use inside New York City-based NYU Langone Health’s HIPAA-compliant GPT instance.
The frontier models won all three, per the June 12 study. On MedQA, Gemini scored 97.4% accuracy versus 89.6% for OpenEvidence and 88.4% for UpToDate. On the real-query benchmark — reviewed blind by 12 clinicians generating 1,800 annotations — the three frontier models formed the top tier, while the clinical tools scored no better than Google’s auto-enabled Search AI Overview.
UpToDate’s AI refused 19% of queries, more than any other model. OpenEvidence scored lowest on clarity, which the authors attributed to communication rather than knowledge gaps. None of the models produced more harmful content or hallucinations than the others.
Senior author Eric Oermann, MD, and colleagues argue the results carry implications for procurement, reimbursement and regulatory oversight, and call for independent, real-world evaluation before clinical AI tools enter practice.
OpenEvidence has asked Nature Medicine to retract the study. In a June 15 letter to the journal, obtained by Becker’s, the company also sought a public apology and an independent review, arguing the paper relies on flawed methods.
The company said the study is vulnerable to “contamination effects” because the publicly available MedQA and HealthBench benchmarks may have appeared in the frontier models’ training data. Peer reviewers raised the same concern, and the authors acknowledged in the paper that the models “may have been exposed to MedQA or HealthBench during training.” The authors designated the third benchmark, built from real clinician queries and free from that risk, as the study’s primary evidence; the frontier models led on that measure as well.
“We continue to stand by our study and its results,” Dr. Oermann told Becker’s.
OpenEvidence also pointed to other evaluations, including a Rochester, Minn.-based Mayo Clinic study, that it said found the tool accurate and guideline-concordant.
Wolters Kluwer also disputed the findings. Peter Bonis, MD, chief medical officer of Wolters Kluwer Health, told Becker’s the study “confused clinical quality and complete-sounding answers” and relied on benchmarks that measure test-taking rather than evidence appraisal or safe use in clinical workflows.
Dr. Bonis pushed back specifically on the refusal rate, saying the paper treated UpToDate Expert AI’s declining to answer as a deficit without establishing whether those refusals were appropriate. In clinical decision support, he said, declining to answer an underspecified or risky prompt “may be the safer behavior.”
He added that UpToDate Expert AI was recently tested on 1,669 clinical queries spanning more than 15,000 criteria and, the company said, returned clinically aligned information for 99.9% of assessed criteria.
Nature Medicine declined to address the retraction demand directly. In a statement to Becker’s, chief editor João Monteiro, MD, PhD, said the journal “welcomes scientific debate” and pointed to its formal Matters Arising process for submitting rebuttals to published papers, which he said allows concerns to be “formally assessed through the journal’s editorial process” alongside a response from the original authors.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.