The healthcare AI PR wars, explained

Advertisement

Clinical AI vendors have spent the past few weeks citing dueling benchmark studies to make the case that their tool is the safest or most trusted by physicians — sometimes citing the same study to reach opposite conclusions.

Here are 12 things to know about the fight, and why the health system leaders who actually decide which tools reach clinicians say it hasn’t changed much for them:

1. The latest flashpoint is NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark study from ARISE, a clinical AI research team of Stanford University School of Medicine and Harvard Medical School physicians. Posted to arXiv in December 2025 and last revised July 13, the study’s primary analysis tested 20 generalist large language models and four clinical decision-support tools against 1,100 case-based tasks drawn from real physician-to-specialist consultations, scored by a panel of 29 board-certified physicians.

2. Ranked by rate of potentially severe harmful errors, four medical-specialized tools tied for the top spot: AMBOSS LiSA (2.9%), Doximity Ask (4.8%), OpenEvidence (5.1%) and Glass Health (5.4%). All four significantly outperformed every general-purpose model tested, including GPT-5.5, Gemini 2.5 Pro and multiple Claude models, as Becker’s reported.

3. Doximity put out a press release July 15 about NOHARM, pointing to a different slice of the data. The company said Doximity Ask ranked first on the study’s “real-world clinical sample,” which it described as the portion of the benchmark that most closely mirrors how physicians use these tools in practice. “We have long believed that the path to trustworthy healthcare AI runs through physicians, not around them,” stated Louis-Antoine Mullie, Doximity’s head of medical AI.

4. OpenEvidence responded July 20 with its own NOHARM release — built around yet another part of the study. OpenEvidence highlighted a companion trial of 101 physicians in which one study arm let clinicians pick any AI tool freely while working real cases. Physicians selected OpenEvidence in 22.3% of their responses, the company said, versus a combined 19.8% for ChatGPT, Claude, Gemini and every other outside chatbot.

5. “The most important benchmark in clinical AI isn’t administered by researchers — it’s administered by physicians, hundreds of thousands of times a day, every time they decide what to consult before making a decision that affects a patient,” OpenEvidence founder and CEO Daniel Nadler said in the release.

6. NOHARM isn’t the only recent clinical AI benchmark to produce dueling narratives. In June, a study published in Nature Medicine found that general-purpose models from OpenAI, Google and Anthropic — GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 — outperformed OpenEvidence and Wolters Kluwer’s UpToDate Expert AI on three benchmarks, including one built from real clinical queries at New York City-based NYU Langone Health.

7. OpenEvidence asked Nature Medicine to retract that study, arguing in a June 15 letter that the benchmarks used were compromised because the models could have seen them during training. Senior author Eric Oermann, MD, told Becker’s: “We continue to stand by our study and its results.” Wolters Kluwer separately disputed the paper’s methodology. Nature Medicine pointed both companies to its formal Matters Arising process for rebuttals rather than acting on the retraction request.

8. Two weeks after that dispute became public, a new preprint flipped the result. Posted to arXiv June 27 and not yet peer-reviewed, the study had 149 physicians across 36 states rate AI answers to real point-of-care questions and found OpenEvidence beat Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 by 25 to 39 percentage points across five dimensions. OpenEvidence codesigned the data collection plan, administered the survey and paid participating physicians, according to the paper, though its authors — based at University of California San Francisco, Harvard Medical School and Stanford University, among others — reported no affiliation with the company.

9. For the CIOs, chief medical information officers and chief AI officers who decide what actually reaches clinicians, the fight has changed little. Rebecca Mishuris, MD, chief health information officer at Somerville, Mass.-based Mass General Brigham, offered a one-word take on the broader dispute: “Noise,” Becker’s reported. At Coral Gables-based Baptist Health South Florida and New York City-based Mount Sinai Health System, leaders said published benchmarks function mainly as an early screen before internal governance committees run their own validation against local patient populations — which determines whether a tool reaches the bedside.

10. Girish Nadkarni, MD, Mount Sinai’s chief AI officer, said the tension between “enormous, deeply resourced” general-purpose AI developers and smaller but still “substantial” domain-focused clinical vendors is real regardless of any single study’s outcome, and expects it to continue. “I genuinely don’t know who ‘wins’ that, and it may not even be the right question, since the answer could differ task by task and shift as both sides evolve,” he told Becker’s. “What I care about is that the way we judge these tools stays honest while it plays out.”

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Register to Attend Webinar

Designing the intelligent hospital: Building hospitals around people, data and care

Tuesday, July 21
12:00 PM - 1:00 PM CDT

Presenters: Braheem Santos, Schneider ElectricJohn Donohue, Penn Medicine

Advertisement

Next Up in Artificial Intelligence

Advertisement