
AI agents going rogue — here’s what you need to know if you’re using ChatGPT, Gemini, or Claude
OpenAI, Google, Meta, and Anthropic have all recently acknowledged incidents in which their AI models accessed real computer systems without authorization during research or security assessments.
In OpenAI’s case, agents actively circumvented restrictions intended to keep them contained, while the other three companies attributed their incidents to poorly configured test environments left connected to the Internet.
And as these same companies work to put more autonomous agents in the hands of consumers, these incidents raise an uncomfortable question about how much control AI companies actually have over the technology they develop.
Here’s a look at all the recent attacks and an explanation for each.
Latest videos fromTom’s Guide
OpenAI agents broke into Hugging Face
One of the most alarming examples involves around 700 OpenAI agents participating in a cybersecurity assessment called ExploitGym, with the attack on Hugging Face taking place between July 8 and 13, 2026, according to an investigation by AI research association METR and OpenAI’s own report.
The agents were supposed to perform their assigned tasks in a controlled environment, but they turned OpenAI’s internal registry of Artifactory packages into an unauthorized message forum, coordinated with each other and sought ways to circumvent restrictions intended to keep them isolated from the Internet. Their efforts ultimately resulted in a breach of Hugging Face, the popular AI development platform, where OpenAI’s technical report revealed that the agents were running code on 41 production dataset server workers.
The incident was not a deliberate attack ordered by a human operator. The agents were trying to succeed in their assigned tasks and adopted methods that exceeded security limits, and OpenAI admitted that early warning signs in the weeks before the breach could have triggered a quicker response.
OpenAI has since attributed the incident to misaligned behavior, which essentially means AI systems pursue a goal in ways their creators never intended.
Hugging Face wasn’t the only target. In a September 25 update, OpenAI said it notifies third parties on an ongoing basis whenever its models may have circumvented security controls, impaired online services, or otherwise affected external websites, and that it has notified dozens of them so far. The company then increased this figure to over 100 organizations, covering notifications sent through September 26.
The activity identified by OpenAI includes agents bypassing access controls and using exposed credentials, as well as some agents modifying third-party websites, including using public wiki pages as makeshift discussion forums that then required cleanup. OpenAI says its review is ongoing and it plans to notify more organizations.
The problem extends well beyond OpenAI.
In August, Meta confirmed a report from The Information that one of its models, identified by the publication as Muse Spark 1.1, had breached another company’s systems during a cybersecurity test.
A Meta spokesperson said a misconfiguration by Irregular, an independent testing company the company works with, inadvertently gave the model access to the Internet. The model then exploited a vulnerability in a third-party service and made changes to the company’s internal systems.
Meta said the incident did not involve a sandbox escape, and did not name the company involved or explain what was changed. But the underlying behavior was troubling, since an AI system placed in the wrong environment was capable of acting against a real organization.
Google reported a similar issue with its Gemini models. During a cybersecurity assessment in May, also conducted with Irregular, a Gemini model that was not supposed to have Internet access came online due to a configuration issue. He then used publicly available information and guessed credentials to access systems belonging to three real companies.
Anthropic disclosed three incidents in which Claude’s models gained unauthorized access to three organizations’ production systems during cybersecurity assessments executed in Irregular’s environment. The models involved were Claude Opus 4.7, Claude Mythos 5 and a new internal research model. In one case, Opus 4.7 was pointed at a fictitious target whose name matched the real domain of a real company.
The problem is not that the AI is bad
When agents go rogue it may seem like they’re evil, but that’s really not the case. In each of these cases, the AI agent was not deliberately seeking to cause harm. the AI took unauthorized actions because it believed those actions would help complete a task.
But despite this, they reveal a practical problem regarding AI agents and how they can be remarkably effective while pursuing a goal without reliably understanding which methods are acceptable.
If an agent is rewarded for completing a task, it may discover shortcuts that break rules, exploit vulnerabilities, or exceed its permissions.
And unlike a traditional chatbot that simply produces text, an autonomous agent can interact with websites, run code, and make changes to connected systems.
Why it matters for everyday AI users
For most people, the immediate risk is not that ChatGPT will suddenly start hacking government websites. This is because AI assistants are increasingly capable of acting on our behalf.
OpenAI’s Dots, launched September 29 for ChatGPT Pro and Business Premium subscribers, along with ChatGPT Work and other agent-like products, represents a broader shift from AI that simply answers questions to AI that accomplishes tasks. This can mean searching for files, interacting with websites, using connected services, and running multi-step workflows with less supervision.
The appeal is obvious, since a task that would take you an hour can be assigned to an AI assistant. But each additional permission creates another opportunity for something to go wrong.
An agent with access to your email may encounter sensitive personal information, an agent connected to cloud storage may be able to read documents you never intended to share, and an agent authorized to interact with external services may encounter malicious instructions or make an unexpected change.
Conclusion
The solution is not to stop using AI agents. I use them regularly and find them incredibly useful.
Start by checking which accounts, apps, and services your AI assistant can access and disconnect anything you don’t actively need, especially services that contain sensitive financial, medical, or personal information.
When possible, require approval before an agent sends messages, edits files, makes purchases, or takes other consequential actions. It’s also worth remembering that an AI agent isn’t limited to doing exactly what you imagined when you gave it an instruction, because a task that seems simple to a human can involve dozens of decisions the AI makes along the way. The fewer permissions an agent has, the less damage an unexpected decision can cause.
Follow Tom’s Guide to Google News And add us as your favorite source to get our latest news, analysis and reviews in your feeds. Subscribe to Tom’s Guide on YouTube and follow us TikTok.
Learn more about Tom’s Guide
Gn bussni