OpenAI shelves GPT-6.1 Astra after deceptive behavior surfaces

Advertisement

OpenAI has canceled the public release of GPT-6.1 Astra, its next-generation AI model, after internal testing found safety and alignment problems, the company confirmed to Becker’s.

The model had been slated to debut in ChatGPT and Codex in October. It outperformed earlier OpenAI models at writing and at completing complex tasks end to end without human help.

Saachi Jain, OpenAI’s head of safety systems, told The Wall Street Journal, which first reported the news Sept. 28, that the model performed worse than its predecessor, GPT-6 Astra, in two areas. It showed higher levels of deception and did not always tell users accurately which actions it had or had not taken. It also overstepped user authorization, pushing ahead on tasks without asking permission and sometimes reaching for external tools and services even when that might be unsafe.

In a statement to Becker’s, Ms. Jain said that while GPT-6.1 Astra reduced model laziness, it “didn’t quite meet the bar in terms of staying within scope and authorization.”

OpenAI said it has other new models coming soon that meet its safety bar and plans to release future Astra models. The company plans to reuse the same base model for additional reinforcement learning runs toward future GPT-6 models. It will also investigate the root causes, including whether its reinforcement learning environments reward the right behaviors.

The decision is separate from OpenAI’s pause last week on training its most capable models, which followed an AI agent slipping through a gap in the company’s internet restrictions to query a public chatbot, according to the Journal.

The announcement came one day before OpenAI’s annual developer conference in San Francisco and days before a Senate subcommittee hearing on securing the U.S. against AI agent attacks.

Advertisement

Next Up in Artificial Intelligence

Advertisement