A June study in Nature Medicine found that general-purpose models from OpenAI, Google and Anthropic outperformed specialized clinical AI tools from OpenEvidence and Wolters Kluwer across a series of medical benchmarks. OpenEvidence has since asked the journal to retract the paper, and Wolters Kluwer has disputed its methodology.
For the executives who actually decide which AI tools reach clinicians, the dispute hasn’t changed much — because they didn’t heavily rely on published benchmarks to begin with.
Somerville, Mass.-based Mass General Brigham’s evaluation depends on what outcome a tool is meant to achieve, alongside “balancing measures,” said Rebecca Mishuris, MD, vice president and chief health information officer.
“We look at aspects related to clinical safety and quality, user experience, patient experience — this is in addition to our usual evaluation of privacy, security, model explainability and bias,” she told Becker’s. “We always lean on our responsible use of AI framework, though.”
However, Dr. Mishuris noted that independent studies and Mass General Brigham’s own internal testing carry more weight than a vendor’s validation data. “We want to see how it performs in our environment,” she noted.
At Coral Gables-based Baptist Health South Florida, published benchmarks and vendor materials such as model cards can get an AI product in the door for review but not to the bedside, said Chief Digital and Information Officer Sha Edathumparampil.
“Before any solution is deployed to production, internal AI governance runs a 360-degree review with experts from essentially every department across the health system, including doctors and nurses, followed by pilots that test the tool against real-world data,” he said.
That step matters, Mr. Edathumparampil noted, because vendor models trained and validated on large national datasets can behave differently once they encounter a health system’s specific patient population. “So local validation and a human in the loop are necessary in clinical settings,” he said.
Meanwhile, at New York City-based Mount Sinai Health System, that local validation runs through a dedicated assurance lab built specifically to test AI tools — whether vendor-supplied or developed in-house — before reaching clinicians.
“Public benchmarks are part of the picture, and we pay attention to them, but they really function as an early screen,” said Ben Glicksberg, PhD, chief of innovation and entrepreneurship at Mount Sinai. “A good score tells us a tool is worth a closer look. It doesn’t tell us much about how the thing will actually perform on our patients, answering the questions our practitioners actually ask.”
Most of the lab’s effort, Dr. Glicksberg said, goes into testing on Mount Sinai’s large and diverse patient population using internal benchmarks built around the system’s own use cases — with particular attention to whether performance holds up across subgroups rather than just on average.
Girish Nadkarni, MD, Mount Sinai’s chief AI officer, said the lab spends significant time examining where a tool fails rather than simply how often it succeeds, since the cost of an error isn’t evenly distributed.
“One confidently wrong answer can matter more than numerous right ones,” he said. “Published benchmarks help us decide what’s worth bringing in for that kind of scrutiny, but they were never going to clear something for the bedside on their own.”
One flashpoint in the dispute was less clear-cut for health system leaders: how to interpret a model’s refusal to answer.
Dr. Mishuris said a refusal can be either a safety mechanism or evidence of a gap in training, depending on context.
Mr. Edathumparampil said the more useful question isn’t whether a tool refuses, but whether it does for the right reasons and explains why.
“Taking every refusal as a deficit may encourage the wrong behavior: a model that answers confidently when it shouldn’t, which in a clinical setting is a bad thing,” he said. “But at the same time, ‘declining is always safer’ isn’t great either, because refusal may mask capability gaps.”
Said Dr. Nadkarni: “The part that genuinely worries me is how easily people learn to work around a refusal, such as rewording the question or stripping out the detail that triggered the caution, and suddenly you get the answer the tool was trying not to give. A refusal that’s one rephrasing away from being defeated isn’t really a safeguard.”
Mount Sinai, for one, treats evaluation as ongoing, since patient populations shift and models can drift in performance long after go-live, Dr. Glicksberg noted.
That posture mirrors how other systems have approached AI oversight. New York City-based NYU Langone Health has retired tools that no longer validate as expected after deployment, while Cleveland Clinic and Los Angeles-based Cedars-Sinai have created internal validation tools of their own because vendors often can’t confirm performance against a system’s specific data. Chesterfield, Mo.-based Mercy and Houston-based Texas Children’s have likewise built multidisciplinary governance frameworks that vet tools before they reach patients.
As for the public fight itself, the leaders who weighed in were largely unmoved. Dr. Mishuris’ assessment was one word: “Noise.”
Dr. Nadkarni said the underlying tension is real even if the specific dispute isn’t decisive — a structural pull between “enormous, deeply resourced” general-purpose AI developers and smaller but still “substantial” domain-focused clinical vendors that he expects to continue.
“I genuinely don’t know who ‘wins’ that, and it may not even be the right question, since the answer could differ task by task and shift as both sides evolve,” he said. “What I care about is that the way we judge these tools stays honest while it plays out.”
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.