24 AI tools ranked for patient safety: Stanford, Harvard study

Advertisement

Stanford University School of Medicine and Harvard Medical School researchers ranked 24 AI systems by clinical safety in a new benchmark study, finding four medical-specialized tools topped the list in a statistical tie.

The findings come from NOHARM (Numerous Options Harm Assessment for Risk in Medicine), originally posted to arXiv in December 2025 and last revised July 13.

Researchers tested 20 generalist large language models and four retrieval-augmented generation clinical AI tools against 1,100 case-based tasks drawn from real physician-to-specialist consultations, scored by a panel of 29 board-certified physicians. Below is a ranking by rate of potentially severe harmful errors, from lowest to highest. A hypothetical “do nothing” model that recommended no action in any case scored worst overall, with potential for severe harm in 37% of cases.

Note: this list includes ties.

1 (tie). AMBOSS LiSA: 2.9%

1 (tie). Doximity Ask: 4.8%

1 (tie). OpenEvidence: 5.1%

1 (tie). Glass Health: 5.4%

5. GPT-5.5: 8.9%

6. GPT-5.6 Sol: 9.0%

7. GPT-5.4: 9.8%

8. Gemini 2.5 Pro: 12.8%

9. Claude Fable 5: 13.6%

10. MedGemma 27B: 14.4%

11 (tie). Qwen3.5 397B: 14.8%

11 (tie). Claude Opus 4.7: 14.8%

13 (tie). Claude Sonnet 5: 15.3%

13 (tie). Kimi K2.5: 15.3%

15 (tie). DeepSeek R1: 15.4%

15 (tie). Claude Opus 4.8: 15.4%

15 (tie). Kimi K2.6: 15.4%

18. Gemini 3.1 Pro: 15.6%

19. GLM 5.1: 16.4%

20. Claude Sonnet 4.6: 16.6%

21. DeepSeek V4 Pro: 19.0%

22. Grok 4.3: 20.9%

23. MedGemma 1.5 4B: 21.4%

24. Llama 4 Maverick: 24.6%

The four top-ranked clinical AI tools showed no statistically significant differences among themselves, but each significantly outperformed every generalist model tested. Errors of omission, meaning the AI recommended too little rather than something harmful, accounted for more than 80% of severe errors across all systems.

A companion randomized trial of 101 U.S. physicians found AI assistance improved their performance over conventional resources, but physicians who used AI still scored well below the top-performing standalone AI systems, often because they skipped valuable recommendations the AI had already surfaced.

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Register to Attend Webinar

Designing the intelligent hospital: Building hospitals around people, data and care

Tuesday, July 21
12:00 PM - 1:00 PM CDT

Presenters: Braheem Santos, Schneider ElectricJohn Donohue, Penn Medicine

Advertisement

Next Up in Artificial Intelligence

Advertisement