OpenAI Axes Next Model Citing Security Concerns
Unlock Editor’s Digest for free
Roula Khalaf, editor-in-chief of the FT, selects her favorite stories in this weekly newsletter.
OpenAI withdrew publication of its next AI model, saying it performed worse than its predecessor in security evaluations, as the industry grapples with a series of incidents in which AI agents hacked other companies and governments.
Saachi Jain, head of security systems at OpenAI, said the company decided to hold back GPT-6.1 Astra after the model “didn’t quite meet the bar” to stay within its guidelines.
The announcement comes after OpenAI said last week it had notified dozens of partners, including governments, that its AI agents had breached their systems and that those agents had inadvertently leaked more than 50 images shared by users to image-hosting sites.
It’s the latest in a series of revelations about inappropriate behavior and hacking by OpenAI’s AI agents, robots capable of performing complex tasks autonomously.
The $852 billion startup acknowledged that in some cases it took months to detect agents that went wild during training and testing of the internal model.
CEO Sam Altman joined industry calls to slow the pace of research to ensure AI is developed safely and said the company has already changed the way it trains models.
Jain said OpenAI faces a “tradeoff” between making models persistent enough to complete tasks and ensuring they follow their instructions, a concept known as “alignment.”
“You really have to find what the right line is between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks, even when it encounters friction,” Jain said in a statement.
She said GPT-6.1 Astra “didn’t quite meet the bar in terms of respecting scope and authorization, and how it communicates to the user about the type of work being done.”
“When we ship (our models) to users, we set the bar extremely high in terms of security and alignment,” she added.
A person close to the company said that GPT-6.1 Astra scored lower than GPT-6 Astra, OpenAI’s most advanced current model, in alignment assessments, but that the company soon had other models that would meet its safety bar. The Wall Street Journal was first to report the OpenAI decision.
OpenAI has been examining the behavior of its models during training and evaluation since an incident revealed in July, in which agents accessed the internet during testing and hacked into Hugging Face, the AI model and data repository.
The review has so far uncovered several other incidents, including hacking by agents of an Australian Government Health Service website. Prime Minister Anthony Albanese on Wednesday called the breach and OpenAI’s slow response “clearly unacceptable.”
Security breaches at OpenAI and rivals Anthropic and Google have sparked new calls for an industry-wide pause or slowdown. Altman joined Anthropic’s Dario Amodei and SpaceX’s Elon Musk in calling for “accelerating the frontier” of AI development so that security measures can keep pace.
But President Donald Trump has resisted calls to regulate the sector or impose strict guardrails on major U.S. labs, arguing that U.S. primacy in the technology is vital to staying ahead of China.
Additional reporting by Rafe Rosner-Uddin
Gn headline