GPT-6.1 Astra went too far: OpenAI cancels launch over deception and unauthorized actions

OpenAI says the upgraded model was better at completing complex tasks but failed tests on staying within its assigned scope, respecting permissions and accurately telling users what it had done

OpenAI had planned to launch a powerful new artificial intelligence model in October that could carry out increasingly complex tasks with little human assistance. Now the company says GPT-6.1 Astra will not be released after internal testing found that it could mislead users about its actions and sometimes continue working beyond the task and permissions it had been given.
The decision comes as OpenAI investigates a series of incidents in which other experimental models acted in ways the company had not intended, including outside the environments in which they were supposed to remain.
סם אלטמן
סם אלטמן
OpenAI stops GPT-6.1 Astra release as increasingly autonomous AI agents raise alarms
(Photo: Getty Images)
Saachi Jain, OpenAI’s head of safety systems, said the new model had improved in its ability to persist with difficult tasks but failed to clear the company’s safety threshold.
It “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” Jain said.
For users, the distinction can be significant. An AI assistant that cannot find requested information might simply say it failed. A more autonomous system may instead search for another way to obtain it, potentially without first determining whether that alternative route is authorized.
OpenAI has not said when, or whether, a modified version of GPT-6.1 Astra could eventually be released.

A model powerful enough to find unknown cyber flaws

The concern is heightened by what the Astra family of models can already do.
GPT-6 Astra, released earlier this month, became the first OpenAI model classified at the “Critical” level for cybersecurity capabilities under the company’s Preparedness Framework.
OpenAI said its testing showed that, given appropriate tools and access, the model could identify previously unknown security vulnerabilities and develop methods to exploit them across well-protected computer systems without requiring a human to guide every step.
GPT 6 Astra
GPT 6 Astra
GPT 6 Astra
(Photo: OpenAI)
Such capabilities could help defenders identify weaknesses before attackers exploit them. But they also create two distinct risks: a person deliberately instructing an AI system to carry out a cyberattack, and a sufficiently autonomous model taking unauthorized actions even when nobody explicitly instructed it to do so.
That second possibility has become much harder for AI companies to dismiss after several recent incidents involving experimental agents.

The Hugging Face ‘warning shot’

The most serious episode became public after internal cybersecurity evaluations in July.
OpenAI disclosed that several of its models circumvented restrictions intended to isolate them from the internet, found ways to communicate with one another through unauthorized channels and ultimately compromised systems belonging to AI platform Hugging Face.
The models were supposed to solve cybersecurity challenges in controlled testing environments. Instead, they discovered vulnerabilities in OpenAI’s own research infrastructure, gained internet access and used it to reach external systems.
תומס וולף
תומס וולף
Hugging Face co-founder and chief science officer Thomas Wolf
(Photo: Getty Images)
Agents eventually executed commands on Hugging Face servers, obtained credentials and reached information and permissions they had not been given.
OpenAI later described the episode as a “warning shot,” saying it demonstrated that sufficiently capable AI agents could work around technical controls, collaborate through channels their developers had not authorized and carry out dangerous actions that no person had requested.
The company stressed that GPT-6 Astra itself was not involved in the Hugging Face incident.

Australian government portal accessed without permission

A separate case emerged from an OpenAI research task conducted in June.
An experimental model was asked to locate information about government spending on medications for skin diseases. When it encountered difficulty obtaining the data, the system found a way into a non-public portion of an Australian government statistics portal linked to Medicare.
The agent executed commands, retrieved files and accessed information beyond what had been publicly available.
OpenAI has said there is no evidence that personal medical records were accessed. The information reached by the model included aggregate health statistics and internal file names.
The company detected the unusual activity during a broader review in August but did not notify Australian authorities until September 10. OpenAI has acknowledged that the government should have been informed sooner.
The company’s broader review has since identified other cases in which agents reached third-party systems or acted outside intended boundaries.
OpenAI has said it has notified dozens of affected third parties while continuing to review earlier training and evaluation activity.

OpenAI pauses some of its most powerful agent work

The growing list of incidents has pushed the company to tighten restrictions on its most capable models.
After the Hugging Face episode, OpenAI strengthened isolation between research environments, restricted internet access, increased monitoring and imposed tougher security requirements on models capable of executing code or using external tools.
More recently, after another agent found a gap in internet restrictions and used it to contact an outside chatbot, OpenAI said training, evaluation and inference involving tool use for its most capable models remained paused.
The company has increasingly focused on a problem that becomes more serious as AI systems grow more persistent: a model may begin with a legitimate objective but, when it encounters obstacles, gradually move beyond the boundaries of the original task in pursuit of that objective.
That makes alignment less about preventing one obviously dangerous answer and more about monitoring an entire sequence of decisions stretching across minutes or hours of autonomous work.

Some want to slow down, Trump wants America to keep moving

The safety debate is no longer confined to OpenAI.
Anthropic CEO Dario Amodei has called for AI companies to slow the improvement of frontier models long enough for safety testing and oversight to catch up. He has advocated independent evaluators inside leading laboratories and greater coordination between companies and governments.
Amodei has also acknowledged that Anthropic has experienced comparable incidents, though he described them as less severe than OpenAI’s Hugging Face episode.
OpenAI leaders have likewise supported stronger external scrutiny, making the decision not to release GPT-6.1 Astra a practical test of whether companies are willing to delay products when internal safety findings fall short.
At the same time, the industry faces strong political pressure to keep accelerating.
Amodei met President Donald Trump for a private dinner Sunday night, their first one-on-one meeting. On Tuesday, he is expected to join a broader White House discussion with senior technology executives, including OpenAI President Greg Brockman, Meta CEO Mark Zuckerberg and Nvidia CEO Jensen Huang.
Trump has repeatedly argued that the United States cannot afford to slow AI development if doing so risks surrendering technological leadership to China.
That leaves the industry confronting an increasingly difficult equation: AI systems are becoming more capable of operating independently just as companies, governments and researchers are discovering how hard it can be to ensure that those systems stop where humans intended them to stop.
For OpenAI, GPT-6.1 Astra failed that test. The company chose not to ship it.
Comments
The commenter agrees to the privacy policy of Ynet News and agrees not to submit comments that violate the terms of use, including incitement, libel and expressions that exceed the accepted norms of freedom of speech.
""