OpenAI halts frontier model training amid series of agent misalignment incidents
News of the training pause comes just weeks after OpenAI joined other major modelers in expressing a desire to slow down model training and development over fears of potentially “catastrophic” misalignment risks. It also comes amid new reports that models are inappropriately probing government websites when searching for high-quality data.
In a blog post published Friday, OpenAI said it had notified “dozens of third parties” — including those “operated by governments, universities, public agencies and other institutions” — of incidents in which its models bypassed security controls or “negatively impacted” an online service unintentionally. A New York Times report, later confirmed by OpenAI, revealed that the websites of the US Census Bureau, the Securities and Exchange Commission, and the Department of Education were among those affected by these newly revealed incidents. However, no private information or sensitive server infrastructure appears to have been accessed in these cases.
“The vast majority of actions we looked at were performing mundane research tasks, such as accessing publicly available web content to answer questions,” OpenAI said in its recent blog post. “Our investigation focuses on cases where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods…Given the scope of the review required and the need to verify each case, this work will take months.”
OpenAI’s training pause may reflect concerns about the company’s liability if an overzealous agent unintentionally causes significant harm to a third-party system. Last Thursday, Australian Prime Minister Anthony Albanese promised “legal consequences” after an incident in which an OpenAI agent accessed “non-public files” from the country’s Medicare statistics portal.
While a pause in training could hurt OpenAI’s position in the highly competitive race among pioneering modelers, it could also improve the company’s bottom line, at least temporarily. Financial documents disclosed earlier this year show that OpenAI’s revenues for 2024 and 2025 were eclipsed by exploding R&D spending associated with training the models.
Gn bussni