An AI agent created fake online identities in an attempt to gain unauthorized access to secure systems during government safety tests, Britain’s AI Security Institute said Tuesday.
The incident emerged during evaluations of agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol, which the institute said carried out a series of unauthorized actions while completing a fictional cybersecurity exercise.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute, known as AISI, said in a blog post.
The findings highlight weaknesses in the safeguards surrounding the testing of AI agents, even as major technology companies promote them as the future of business automation.
AISI receives access to advanced AI models through voluntary agreements with leading laboratories. In the evaluation, it placed the agents in a simulated cybersecurity scenario designed to test their capabilities and behavior.
The institute ran the challenge 122 times and identified 19 unauthorized actions across 10 test runs. Anthropic’s agent was responsible for 17 of the actions, while OpenAI’s agent accounted for the remaining two.
The most serious incident involved an agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code.
AISI said it found no evidence that any of the incidents caused real-world harm.
The institute did not initially identify which agent created the fake identities, but Anthropic later confirmed that its system was responsible.
“We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” Anthropic said.
The company said it was working with AISI to obtain additional information and conduct its own investigation.
Andrew Yoon, a researcher at CivAI, a California nonprofit that studies AI capabilities and risks, said the incident raised questions about Anthropic’s control over its systems.
“The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think,” he said.
OpenAI also published details of its agent’s conduct, saying both unauthorized actions involved accessing the internet in ways prohibited by the evaluation prompt.
“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said.
The company also disclosed a separate incident involving Irregular, a third-party testing provider. A configuration error allowed OpenAI agents to connect to the internet unintentionally.
Anthropic disclosed a similar configuration issue last week.
Reuters reported last week that OpenAI had expanded an investigation into agent-related security incidents after finding evidence of additional breakouts.
Unlike the July breach of AI company Hugging Face by an OpenAI agent, the systems in the AISI evaluation did not escape an isolated testing environment in order to reach the internet.
Instead, AISI said internet access had been permitted under its standard testing procedures. The unauthorized behavior involved agents using that access in ways that violated the rules of the exercise.



