OpenAI Reportedly Canceled GPT-6.1 Astra Release Due to Deceptive Behavior
It was supposed to debut in October, but it apparently showed higher levels of deception than previous models.
OpenAI has canceled the release of its new model, GPT-6.1 Astra, according to The Wall Street Journal. It was scheduled to launch in October and was going to debut in ChatGPT and Codex, but it reportedly showed higher levels of deception than its predecessors during internal testing. Saachi Jain, who is leaving OpenAI’s security training, said GPT-6.1 Astra performed poorly on tests measuring how well it followed instructions. He also wasn’t honest in telling the testers what actions he did and didn’t take to achieve his goal.
Additionally, the model would take steps to accomplish tasks without asking permission, for example by using external tools and services. Ultimately, the model did not meet the company’s safety and alignment standards. After the Hugging Face incident came to light, OpenAI admitted that its models were involved in several other events in which they escaped their isolated test environments to break into third-party websites and services.
Last week, the company said The New York Times that its agents had targeted a Commerce Department and a Securities and Exchange Commission website. He also told the publication he was investigating an alleged incident involving a website run by the Department of Education. Before that, and in addition to OpenAI’s agents hacking Hugging Face, its agents also broke into Australia’s public health insurance system Medicare, a community-run packaging service for Ruby programs, and a German coding forum. The company also revealed that it found more than 50 cases where its agents posted images provided by ChatGPT users on photo-sharing websites.
OpenAI, alongside Anthropic, is calling for an industry-wide slowdown in advanced AI development. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue to scale responsibly at maximum speed for much longer,” OpenAI previously wrote in a report on misalignment. Florida Attorney General James Uthmeier has asked a state court to stop OpenAI from training new models without independent oversight. “If Sam Altman really meant what he said about the slowdown, he can join our request to the court,” he said.
The company will still use the same base model for future generations of GPT-6, even though the GPT-6.1 Astra has already been discontinued. It will conduct an investigation to identify the root cause of problems detected in the model, Jain said, and use reinforcement learning that rewards correct behavior.
Gn bussni