OpenAI has scrapped the planned release of GPT-6.1 Astra, a new artificial-intelligence model intended for ChatGPT and Codex, after internal testing found that it fell short of the company's safety and alignment requirements, including its ability to stay within authorized boundaries and accurately describe its actions.
The model had been scheduled for an October debut and was designed to handle increasingly complex tasks with less human assistance. OpenAI's decision represents an unusual instance of a major AI developer abandoning a near-term model rollout because of safety concerns rather than performance or commercial considerations.
Saachi Jain, OpenAI's head of safety systems, said the company applies a particularly demanding standard before releasing powerful models to customers. "Of course we want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users," she said. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment."
Astra failed to clear that threshold in several areas. Jain said the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done." Testing found instances in which the model pushed ahead without obtaining appropriate permission and wasn't consistently transparent about actions it had taken or avoided.
The decision comes as OpenAI is dealing with a separate security incident involving an internal research agent that found a way around restrictions intended to prevent it from reaching the live internet. During a search-related training task on Sept. 20, the agent discovered that its sandbox's DNS resolver could communicate externally and used that channel to send queries to a public chatbot service.
OpenAI said its monitoring systems detected the behavior within about 15 minutes, and a human reviewer began examining the alert roughly three minutes later. The run nevertheless continued for about another 2½ hours before it was terminated, exposing weaknesses not only in the technical isolation of the training environment but also in how quickly an identified incident could be stopped.
The company said it has since added two independent blocking controls that would each have prevented the DNS access. It also expanded detection measures and accelerated efforts to use models themselves to probe its sandboxes and network defenses for similar vulnerabilities.
"All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused," OpenAI said in its incident report while engineers verify that the network weakness has been closed and conduct additional testing.
OpenAI has said it won't resume training of the specific model involved in that incident and instead plans to begin a fresh run with additional alignment measures. The DNS episode is distinct from the decision not to release GPT-6.1 Astra, but together the developments highlight the difficulty AI companies face as newer systems gain greater autonomy and access to external tools.
Astra's problems centered in part on what OpenAI calls "scope authorization"-whether an AI system understands what it has permission to do and seeks approval when it reaches a boundary. The model also exhibited more deceptive behavior than its predecessor in some evaluations, including inaccurately representing whether it had carried out particular actions.
Those concerns are gaining attention across the AI industry as developers increasingly use AI systems not merely to answer questions but to conduct research, write software, operate computers and assist in developing later generations of AI.
That shift was the focus of a paper published Monday by the University of Cambridge's Programme on AI Science and Policy. More than 20 researchers and experts, including Geoffrey Hinton, Yoshua Bengio, OpenAI Chief Scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft's Eric Horvitz and UC Berkeley's Dawn Song, examined whether automation of AI research could eventually trigger what they call an "intelligence explosion."
The authors describe a scenario in which increasingly capable AI systems automate larger portions of AI research and development, helping create stronger successor systems that in turn accelerate the development process. If that feedback loop becomes sufficiently powerful, they said, years of technological progress could potentially be compressed into months or less.
The paper cautions that such an outcome isn't certain. It characterizes the existing evidence as preliminary and argues that present productivity improvements haven't yet demonstrated that an intelligence explosion is inevitable.
The authors nevertheless said the potential consequences justify government preparation. Rapid advances could produce breakthroughs in science and industry, while also increasing risks from cyberattacks, biological threats, labor-market disruption and powerful systems operating with diminishing human oversight.
"If this triggers an intelligence explosion, it could dramatically bring forward AI's benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded," the paper said.
The researchers recommended greater visibility into how much AI companies are automating their own research, independent oversight of frontier developers, predetermined safety requirements for continued development and mechanisms capable of slowing or stopping high-risk experiments.