Sponsored

General-Purpose AI Is Winning Medicine’s Easy Benchmarks, Radiology Isn’t One of Them.

Advertisement

We recently ran a set of mock radiology examinations, modelled on the UK Royal College of Radiologists FRCR 2B Short Cases. Against a pass threshold of 73.2, our foundation model Harrison.Rad 1.5 scored a median of 86.5; GPT-5.4, the best general-purpose model, scored 44. Other models all scored below 38.

The size of the gap points to something health-system leaders should understand: intelligence is multidimensional and applying the right benchmarks and tests is critical to picking the best model for the task.

Why the headlines can mislead

You may have seen a recent Nature Medicine paper reporting that general-purpose language models outperform specialized clinical AI tools across several medical benchmarks. The convenient reading is that generalist models beat specialized medical AI, but I believe this to be too broad, particularly in radiology.

Look at what was tested: licensing-style questions, clinician preference alignment, and de-identified clinical queries. These are text in, text out. On those tasks, frontier models did better, which is meaningful. But it doesn’t tell you how these models read a scan.

The signal is in the image, not the text

Radiology is different because the information that matters is not conveyed in text, but in images. The clinical notes may reveal an extensive smoking history, but the cancerous lung nodule itself can only be found in the pixels. A guideline may tell you what to do with a finding, but it does not contain the visual signal you need to detect it.  Frontier models can reason their way into a treatment plan, but reasoning cannot surface a finding it never saw in the image.

To confidently detect these abnormalities requires training upon hospital data, PACS data, radiologist reports and addendums, with the variation of real scanners and protocols. A frontier model can be extraordinarily capable and still have seen very little of the data that matter most for reading a scan. That is why a model trained on radiological imaging behaves so differently on an image-interpretation exam.

Make the test match the task

The deeper issue is one every health system already applies to its own quality programs: does the benchmark actually measure what you care about? A text-based medical benchmark tests knowledge and reasoning in language. A radiology short case is closer to the real work of a radiologist: look at the images, find what is relevant, ignore what is not. Neither is a perfect stand-in for clinical practice, but they measure very different things, and only one of them tells you about image interpretation.

So be wary of impressive scores on the wrong test. A model can ace a medical-text leaderboard and still be the wrong tool for chest X-rays.

Holding ourselves to the same standard

Our 86.5 is an internal result: we have not released the examination set, and Harrison.Rad 1.5 (Rad 1.5) has not yet faced external scrutiny, so it should be treated as a useful stress test, not the final word. That scrutiny is exactly what we welcome external researchers to partner with us on.

In a 2026 American Journal of Roentgenology study, radiologists at Stanford and Mass General Brigham compared four AI systems on 212 chest radiographs. Harrison.Rad 1 (the earlier version of the model) produced the reports radiologists most often accepted (up to 75.5%, versus 16 to 57% for the others), rated highest for quality, and preferred most often. Separately, at the 2025 American College of Radiology Annual Meeting, a challenge run with Mass General Brigham had 113 radiologists blind-rate Harrison.Rad 1 reports acceptable 65.4% of the time, against 79.6% for radiologist-written reports. The independent evidence to date is on Harrison.Rad 1, the earlier version; Harrison.Rad 1.5 will be held to the same standard.

The Nature Medicine paper may be a useful result for medical question-answering and sitting examinations but the practice of radiology is a different task, and if you are choosing AI to read scans, you should insist it be evaluated as one.

Harrison.Rad 1 and Harrison.Rad 1.5 are research-only foundation models and are not medical devices regulatory-cleared for clinical use. Harrison.ai is seeking regulatory clearance for medical devices powered by these models in various markets.

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Advertisement

Next Up in Artificial Intelligence

Advertisement