Radiologists, AI models fooled by deepfake X-rays: Study

Advertisement

AI-generated “deepfake” X-ray images were frequently mistaken for real ones by both radiologists and large language models, raising concerns about cybersecurity and diagnostic reliability.

The retrospective study, led by researchers at the Icahn School of Medicine at Mount Sinai in New York City and published March 24 in Radiology, tested 17 radiologists across 12 centers in six countries on 264 images, half of which were authentic and half of which were AI-generated.

When unaware of the images’ synthetic nature, only 41% of radiologists spontaneously identified fakes. After being informed, their mean accuracy in detecting synthetic X-rays rose to 75%, with individual performance ranging from 58% to 92%. No correlation was found between experience level and detection accuracy, though musculoskeletal specialists performed significantly better.

Four large multimodal language models — GPT-4o and GPT-5 (OpenAI), Gemini 2.5 Pro (Google) and Llama 4 Maverick (Meta) — scored between 52% and 89% accuracy. GPT-4o, which created some of the images, was the most accurate detector among the models.

Researchers emphasized the urgent need for protective measures such as digital watermarks and cryptographic signatures and warned of potential future threats from synthetic 3D imaging.

At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.

Advertisement

Next Up in Radiology

Advertisement