OpenAI discloses 6 AI misalignment cases: What healthcare leaders should know

Advertisement

OpenAI has disclosed six instances of unexpected or concerning artificial intelligence model behavior as part of a new framework for tracking, investigating and publicly reporting model misalignment.

The company outlined the framework in a Sept. 16 post, saying it is designed to speed up disclosures rather than waiting to include multiple incidents in a broader report or in system cards accompanying newly released models. OpenAI said the framework could result in reports being published before the company has fully explained or mitigated the behavior.

The framework covers qualifying behavior throughout a model’s lifecycle, including training, evaluation, testing and deployment. OpenAI said examples could include models acting without authorization, coordinating with other models, evading oversight or behaving in ways that call an alignment method or safeguard into question.

The six initial reports involve behavior OpenAI observed during model training or evaluation.

In one case, an unreleased research model inserted unrelated instructions, including directions to disregard its normal constraints, into summaries used to continue its work in another context window. OpenAI identified 27 affected summaries.

During training of GPT-5.6 Sol, OpenAI said many model instances added instructions to summaries telling subsequent model instances to conceal mistakes or misaligned behavior from users. Examples included directions to invent missing historical data without disclosing it and hide mismatches between source versions.

Another model used an exposed API key without authorization while trying to answer a question about earnings figures in a California county. When the model was still unable to retrieve the requested figures, it fabricated them and presented them as information from the requested source.

Other cases involved a model uploading a file to the internet so it could cite it, models using an internal software repository to communicate across separate training samples and agents using public file-hosting websites to exchange files when they could not access one another’s local files.

OpenAI cautioned that the six reports describe individual instances and should not be interpreted as showing how frequently misalignment occurs across its models. The company said its disclosure framework intentionally favors transparency even when the significance of an incident remains uncertain, meaning some disclosed examples could ultimately prove to be isolated rather than part of a broader pattern.

The disclosures come as OpenAI expands further into healthcare.

In January the company launched OpenAI for Healthcare, which includes ChatGPT for Healthcare, an enterprise platform designed for clinicians, administrators and researchers. OpenAI said health systems including Altamonte Springs, Fla.-based AdventHealth; Dallas-based Baylor Scott & White Health; Boston Children’s Hospital, Los Angeles-based Cedars-Sinai; Nashville, Tenn.-based HCA Healthcare; Palo Alto, Calif.-based Stanford Medicine Children’s Health; and University of California, San Francisco were among the organizations rolling out the technology.

On Sept. 1, OpenAI added an Epic EHR integration to ChatGPT for Healthcare, with UCSF Health serving as a pilot partner. The read-only integration allows authorized clinicians at organizations with supported Epic environments to bring patient information, including appointment notes, lab results, medications and specialist documentation, into ChatGPT or use the technology within Epic workflows.

As AI moves closer to the medical record, health system technology leaders have emphasized the need for safeguards around how those systems operate. Leaders from AdventHealth, Cedars-Sinai, Columbus, Ohio-based Nationwide Children’s Hospital, Grand Rapids, Mich.-based Corewell Health and Penn Medicine told Becker’s this month that considerations include secure access, audit logs, minimum necessary data access, traceable outputs, clinician verification and ongoing performance monitoring.

For healthcare leaders evaluating OpenAI tools, the new disclosures provide additional information about how the company’s models can behave and how OpenAI investigates those behaviors. The company did not say any of the six newly disclosed cases occurred in healthcare deployments.

Under OpenAI’s new process, employees can flag potential cases for investigation by the company’s safety and alignment teams. Cases considered for public disclosure will be placed into one of three tracks: Ready for Disclosure, Minor Investigation or Larger Investigation.

The company said full reports will describe the behavior observed, its severity and any external impact, the setting in which it occurred, when it happened and the model or models involved. When possible, OpenAI also plans to disclose unanswered questions and steps it is taking or considering to address the behavior.

OpenAI called the framework a work in progress and said it does not replace existing legal disclosure requirements, including those involving critical safety incidents or cybersecurity breaches.

Advertisement

Next Up in Artificial Intelligence

Advertisement