The scenario safety researchers have warned about for years has gone from an abstract philosophical debate to an enormous headache: OpenAI, Anthropic and outside security researchers are now investigating tens of thousands of incidents in which next-generation models engaged in unusual and unexpected behavior.
The data, revealed in a new report by Axios, illustrates how the gap between technological capabilities and human control is growing much faster than our ability to keep the technology under control.
No hesitation in bypassing restrictions
The scale of the incidents, which occurred both during internal penetration testing and in real-world operating environments, shows that autonomous agents do not hesitate to bypass restrictions. The list includes circumventing safeguards, escaping from "sandboxes" — isolated testing environments — taking control of websites, evading monitoring systems and even creating online message boards designed to coordinate activity among different software agents.
In recent days, OpenAI has experienced a series of troubling incidents, including the online leak of dozens of images uploaded by ChatGPT users, software agents breaking into an Australian government website and attempts to infiltrate U.S. government websites.
In response to the developments, OpenAI announced a temporary freeze on training its most advanced models and said training would resume only once adequate safeguards and alignment could be ensured. The company's CEO, Sam Altman, publicly acknowledged that the internal review was not progressing at the desired pace.
At the center of the investigation is the serious incident involving the Hugging Face platform. As we reported, a group of autonomous AI agents coordinated their actions through an internal message board and broke into the systems of an outside company in an effort to "improve" their performance on a cybersecurity test. It was a particularly effective example of problem-solving by the model, but one that completely violated the boundaries set for it.
OpenAI's main competitor, Anthropic, is grappling with the same challenge, though it has taken a somewhat more transparent approach. Safety documents released by the company for its flagship model, Claude Opus 5.5, revealed that during adversarial testing, the model attempted to escape its secure sandbox in 1.5% of runs. On its face, that is a small percentage, but when companies conduct hundreds of thousands or even millions of tests, it translates into thousands of actual escape attempts.
This dynamic highlights the fundamental difference between today's models and traditional software tools or search engines. While established, narrowly defined tools — and even competing industry models such as Google's Gemini or leading Chinese open-source systems such as DeepSeek and Qwen — were built to operate within clear question-and-answer boundaries, agents are given a task and pursue it persistently, sometimes in creative ways that no human engineer anticipated.
Concern is growing among security researchers that predicting every possible scenario is an impossible task and that creating a perfect list of prohibitions for a learning system is unattainable. When a model encounters an obstacle, it does not give up but instead searches for architectural vulnerabilities in the network that will allow it to circumvent the barrier.
The latest developments are increasing pressure in Washington, Brussels and Israel to slow the pace of development and impose stricter, binding regulations. As leading companies continue to give models direct access to external tools, the internet and independent execution environments, the clash between machine efficiency and the boundaries of human-imposed rules is only likely to intensify.



