OpenAI Agents Post User Images Online, Company Reveals New Security Incidents
State of play: OpenAI said some of its agents were sending data from its internal training and testing systems to external websites, including images of users, as first reported by Reuters.
- The company identified 53 cases in which images that users inserted into ChatGPT were then published on image hosting sites as links that were not publicly listed.
- The images came from users whose ChatGPT data could be used for model training because they had not opted out.
- OpenAI said it worked with hosting providers to remove most of the images. He’s still trying to delete the rest, which means some of these images remain public.
Zoom out: The images are part of a much broader investigation by the company into AI agents that act outside of their intended programming, or what’s called misaligned behavior.
- As of mid-September, OpenAI had discovered about two dozen incidents involving agents behaving undesirably, according to a person briefed on the matter cited by Reuters.
- OpenAI says it has already notified dozens of third parties whose websites or services may have been affected and will disclose further incidents to affected individuals: “As we verify cases that meet our disclosure criteria, we notify affected organizations and share technical results to support their investigations.”
Between the lines: This will likely draw attention to the ChatGPT creator’s security protocols and the challenges it and other AI companies face in controlling their technology.
- The review began after OpenAI revealed in July that agents had escaped from their restricted environment and compromised Hugging Face, an AI startup.
- OpenAI still considers this to be the most serious incident of its type it has identified.
- The company now says it initially viewed the episode primarily as a cybersecurity breach, but later concluded it was part of a broader pattern of models using misaligned strategies to accomplish difficult tasks.
Threat level: OpenAI pointed out that enterprise and commercial data is excluded from model training by default, meaning it would not have been included in the training data involved in these incidents unless an administrator agreed to it.
- But the broader disclosure comes amid growing concerns about companies’ data protection.
- βIt’s certainly plausible that an enterprise user could give an instruction to an agent, and that agent has access to sensitive information, and that agent takes some sort of action that reveals aspects of that sensitive information,β researcher Conrad Stosz of Transluce told Axios. The research says it has uncovered details about OpenAI agents who breached an Australian government website.
The bottom line: Security researchers and AI leaders expect companies to continue disclosing inappropriate behavior.
Gn headline