OpenEvidence said its new Darwin model has become the first AI system in history to score 100% on MedQA, a medical AI benchmark.
Darwin also topped other major medical AI benchmarks, scoring 72.8% on MedXpertQA, 82.7% on HealthBench Professional and 87.2% on NOHARM, ahead of the next-best models, which OpenEvidence named as Claude Fable 5 and Gemini 3.7.
The Sept. 3 announcement introduced a new family of four medical AI models named for figures in the history of medicine. Three of the models — Osler, Sackett and Snow — are production tools available free to verified clinicians. Osler, the fastest at roughly five seconds per answer, replaces the model that previously powered OpenEvidence and becomes the platform’s default. Sackett takes about 30 seconds per answer for questions that hinge on the weight of clinical evidence. Snow, the successor to OpenEvidence’s Deep Consult feature, runs a full review of medical literature over about five minutes before producing a report.
Darwin, the fourth model, remains in research preview and is available only by application, which OpenEvidence attributed to dual-use risk in fields such as virology, immunology and genetics. The company said institutional partners, including the National Organization for Rare Disorders, and accredited academic researchers currently have access.
OpenEvidence, valued at $12 billion as of January, said it plans to extend Darwin’s capabilities into its production models as safeguards are validated with its partners.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.