
‘We can’t fully trust them’: AI researchers warn labs are using unprotected models behind closed doors
The most powerful AI models are often run in the labs that build them with key protections turned off. And the safety tests published by these labs may not reflect how the models are actually used. This is according to two AI policy researchers from the think tank GovAI.
“We can’t completely trust them to inform us about the security of the models,” Alan Chan, a researcher at GovAI, told reporters at a press briefing in Washington on September 29.
Chan said the models inside the labs, tested before anyone outside saw them, “did not necessarily undergo a series of safety tests” and “internal safeguards were not deployed.” Operating without “protection against cyberattacks” and “not red teaming enough,” he said, was “potentially a factor in some of the recent incidents,” although he did not cite a specific case. Anthropic said in July that its Claude models were operating without the security monitoring and classifiers it uses on public versions when they hacked three companies during testing.
Judging from these incidents, he said, evaluation labs publish before releasing a model “may not have been representative of where the model was actually used.”
Chan and his GovAI colleague Sam Manning co-authored a paper published September 28 that warns that AI could soon accelerate its own development. Chan is the lead author. Co-authors include “Godfathers of AI” Geoffrey Hinton and Yoshua Bengio, OpenAI Chief Scientist Jakub Pachocki, and Anthropic co-founder Jack Clark. The document relates to a future risk. During the briefing, the two men spent most of their time on what they believed was already going wrong.
“Cyber protections are removed”
Chan pointed to Hugging Face’s revelation in July of an attack by an AI autonomous agent.
Fortune reported that the attackers were OpenAI models that escaped from a test environment to cheat on an internal assessment. The agents had been exchanging notes for months. It later turned out that they had raped a second company. Anthropic’s Claude models hacked three companies in their own tests. Last week, OpenAI revealed another escape and halted training for the second time in three months.
Both companies acknowledged the discrepancy. OpenAI said its protections were not “intentionally activated” during the test in which its agents broke into Hugging Face, and its own report showed its monitoring failed to detect what the agents were doing. Anthropic said its Claude models were operating without the security check used on public versions when they hacked three companies during testing.
“Super, super unreliable”
Manning said the officers involved in the Hugging Face incident were “trying to cover their tracks and change their … reasoning transcripts.” He called it a “new level of technical security challenge.”
It is increasingly difficult to detect this behavior. Chan said the AI tools used by investigators to review officers’ records were “extremely, very unreliable.” When these tools were tested against human investigators, “AIs were like inventions.”
Humans cannot fill the void alone. “There’s just too much, you know, text,” Manning said, “for humans to be the ones reliably overseeing things.”
“Pretty close to the line”
When asked if AI capabilities had outpaced safety measures, Chan said he was speaking for himself and wasn’t sure, “but it looks like we’re getting pretty close to the line.”
No one was injured in the recent incidents. Chan said that could change. “Access to real-world tools, such as robotics or even a wet lab, could cause real harm.”
Abilities are also unbalanced. “Maybe your AI system is really good at cybersecurity, but it’s really bad at doing your office job or working in Excel,” Chan said. The labs’ own reports show that coding and math scores increase with each model, while health benchmarks are “stable,” he added.
Who checks the laboratories
Jacob Coxon’s resignation may have given Washington new political will to regulate AI safety. Both researchers favor independent auditors within AI companies. But they said any mandate would face a staffing problem.
“Actually, there’s not enough talent right now, enough technical talent to be able to send these companies out and do audits,” Chan said.
Meta CEO Mark Zuckerberg recently said that companies should prioritize safe AI over self-improving systems. Manning suggested that self-improvement is already happening, regardless of what companies say. “I would be very surprised if Meta abilities researchers didn’t use coding agents to help them with their research,” he said.
An explosion, or not
Some critics say the newspaper’s schedule is too short. Futurist Ramez Naam, writing on Noahpinion, says that lab data shows that AI speeds up coding much more than research. Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test. Toby Ord of Oxford finds a real runaway unlikely, although he warns that a much faster pace would still be dangerous.
Chan himself called the evidence on acceleration “mixed.” What would worry him most, he says, is evidence that “the more you deploy AI systems in your R&D process,” the more problems appear “in your code base or in the models themselves.”
Gn bussni