Researchers at the Icahn School of Medicine at Mount Sinai in New York City found that adding a brief safety reminder to prompts reduced potentially harmful clinical choices in 19 of 20 large language models tested.
The study, published Sept. 26 in Communications Medicine, analyzed more than 10 million AI responses. Researchers tested the models on 501 variations of 50 clinical scenarios, plus 100 cases adapted from deidentified hospital discharge records.
Without a safety reminder, 16.6% of model responses were potentially harmful. With the reminder, that rate fell to 10.1%. Across all responses, the models made about 1.18 million potentially harmful clinical choices.
In one scenario, a model was told to skip recommended follow-up blood tests to reduce workload. Sometimes the request was framed as urgent or as an order from a superior. Other harmful choices included stopping antibiotics early without sufficient clinical reason.
“AI models do not make decisions in a vacuum,” said first author Mahmud Omar, MD, a lecturer in the school’s Windreich Department of Artificial Intelligence and Human Health, in an Oct. 8 news release. Dr. Omar said the reminder did not eliminate harmful choices, so it should serve as one safeguard rather than a replacement for clinical oversight.
Co-senior author Girish Nadkarni, MD, chair of the Windreich department and chief AI officer of Mount Sinai Health System, said safety testing must look beyond whether a model answers correctly under ordinary conditions. As AI systems grow more autonomous, Dr. Nadkarni said, organizations need to know whether they can recognize an unsafe instruction, question it or ask a human for help.
The researchers recommend that developers and health systems build automated safety testing into clinical AI evaluation, both before deployment and after model updates. Next, they plan to study how accumulated context, including prompt injection and time or budget pressures, affects AI agents’ decisions.