OpenAI and Anthropic investigate tens of thousands of security incidents
- The findings, which surface as part of internal model evaluation work and surveys at both companies into model behavior, raise the question of whether either company – or any other major model maker – is currently capable of establishing complete control over its technology.
The details: Episodes include bypassing guardrails, creating chat rooms, escaping sandboxes, hijacking websites, self-prompting or attempting to bypass monitors, sources said.
- They occurred in internal testing and in the real world, and many have not yet been made public as security researchers continue to investigate, sources said.
- Some tests amount to “red-teaming” activities, where companies try to trick models into misbehaving in order to ensure their safety, sources said.
- Inappropriate agent behavior is becoming synonymous with the development of cutting-edge AI: The largest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to accomplish tasks.
Driving the news: Incidents vary in severity and are comparable to disclosures by OpenAI in recent days. These are both successful attempts to circumvent the guardrails and unsuccessful attempts, and so far most have not caused any real damage. The total could well exceed tens of thousands, according to sources.
- In recent days, OpenAI and external researchers have disclosed a litany of episodes involving pattern behavior from the company’s systems that some experts consider troubling.
- These include OpenAI agents leaking 53 online images of ChatGPT users, the breach of an Australian government website and attempts to hack other sites – including those of the US government – according to the company, sources and reports from Reuters and the New York Times.
OpenAI announced it was taking a break training on its highest-performing models and would resume their training “only when we are confident we have additional safeguards and alignment improvements in place,” a spokesperson told Axios.
- Chief executive Sam Altman said on X that the ongoing review had “not been as quick as we would have liked.”
- Altman said the Hugging Face incident was the most serious they had seen. In this case, a swarm of hundreds of agents coordinated their work on a discussion forum and hacked an external company in an attempt to improve their performance in a cybersecurity test.
- “People want to know that AI is developed safely, and that starts with what companies like ours do ourselves,” an OpenAI spokesperson told Axios. “This is not the first time we have paused to take such steps, nor do we think it will be the last as AI capabilities continue to advance.”
State of play: Anthropic has contracted a third-party security organization to review the behavior of its models. In publicly released company documents, the company disclosed the frequency of misalignment episodes.
- The “system map” of its Opus 5.5 model released this week showed how often it behaved in ways the company reported as unusual or problematic.
- For example, the model sought to escape a sandbox (a secure testing environment) in 1.5% of tests, although the company emphasized that these were adversarial experiments in which a task could not be solved without escaping the sandbox.
Yes, but: Anthropic and other companies run hundreds of thousands of tests on their models, if not more, according to sources. This means that even a small percentage of misaligned behaviors can still account for tens of thousands of incidents in which models behaved in unexpected, sometimes troubling, ways.
The Hugging Face Incidentalong with a host of others that followed, led top AI leaders to call for a slowdown in development and push for stricter federal and international regulations.
Some at OpenAI see Hugging Face as one-off, with disclosures about future incidents likely to be less serious due to improved controls and the unusual nature of the tests conducted, which involved a never-before-seen model, sources told Axios.
- AI security researchers agree that there are simple solutions that will help AI companies avoid some aspects of what made the Hugging Face episode so dangerous in the eyes of outsiders.
Threat Level: Other AI executives and security researchers cautioned, however, that they have limited confidence in AI companies’ ability to prevent problematic model behavior.
- The new generation of AI models perform tasks with extraordinary resilience, so working to limit their ingenuity is often a losing game as it is necessary to anticipate all the possible ways they could go crazy.
- Often, a technique that may never have occurred to humans is what allows them to bypass guardrails, senior AI officials said. “Trying to come up with a perfect list of do’s and don’ts is probably a crazy task,” one cybersecurity executive said.
Reality check: What AI security professionals call “misaligned behavior” within AI companies is to be expected as they test their new models.
- Reducing the risk of misalignment to zero may not be feasible, experts told Axios.
Zoom: The problem is that if a model performs a problematic action repeatedly during testing, it is more likely that the model’s behavior will cause a real-world cyber incident.
- “What we’ve seen in terms of what these agents are doing is just the tip of the iceberg,” researcher Conrad Stosz of Transluce, an independent AI evaluator, told Axios.
- It’s not about how damaging each individual instance was, Connor Leahy, an AI researcher and executive director of ControlAI, told Axios.
- What’s “crazy,” he says, is that these cases involve “autonomous systems doing things they’ve been told not to do,” potentially including crimes.
The essentials: Expect more revelations about model misbehavior as AI companies continue to expand their cutting-edge capabilities.
Gn bussni