New security breaches implicate OpenAI and Anthropic AI agents

Britain’s AI Security Institute found 19 unauthorized actions by OpenAI and Anthropic agents, including malicious code, fake identities and prohibited internet access, raising fresh concerns over safeguards for increasingly capable AI systems

An AI agent created fake online identities in an attempt to gain unauthorized access to secure systems during government safety tests, Britain’s AI Security Institute said Tuesday.
The incident emerged during evaluations of agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol, which the institute said carried out a series of unauthorized actions while completing a fictional cybersecurity exercise.
מנכ"ל Open AI, סם אלטמן,
מנכ"ל Open AI, סם אלטמן,
OpenAI CEO Sam Altman
(Photo: TechCrunch)
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute, known as AISI, said in a blog post.
The findings highlight weaknesses in the safeguards surrounding the testing of AI agents, even as major technology companies promote them as the future of business automation.
AISI receives access to advanced AI models through voluntary agreements with leading laboratories. In the evaluation, it placed the agents in a simulated cybersecurity scenario designed to test their capabilities and behavior.
The institute ran the challenge 122 times and identified 19 unauthorized actions across 10 test runs. Anthropic’s agent was responsible for 17 of the actions, while OpenAI’s agent accounted for the remaining two.
The most serious incident involved an agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code.
AISI said it found no evidence that any of the incidents caused real-world harm.
The institute did not initially identify which agent created the fake identities, but Anthropic later confirmed that its system was responsible.
“We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” Anthropic said.
The company said it was working with AISI to obtain additional information and conduct its own investigation.
Andrew Yoon, a researcher at CivAI, a California nonprofit that studies AI capabilities and risks, said the incident raised questions about Anthropic’s control over its systems.
“The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think,” he said.
מנכ"ל אנת'רופיק, דריו אמודיי
מנכ"ל אנת'רופיק, דריו אמודיי
Anthropic CEO Dario Amodei
(Photo: Getty Images)
OpenAI also published details of its agent’s conduct, saying both unauthorized actions involved accessing the internet in ways prohibited by the evaluation prompt.
“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said.
The company also disclosed a separate incident involving Irregular, a third-party testing provider. A configuration error allowed OpenAI agents to connect to the internet unintentionally.
Anthropic disclosed a similar configuration issue last week.
Reuters reported last week that OpenAI had expanded an investigation into agent-related security incidents after finding evidence of additional breakouts.
Unlike the July breach of AI company Hugging Face by an OpenAI agent, the systems in the AISI evaluation did not escape an isolated testing environment in order to reach the internet.
Instead, AISI said internet access had been permitted under its standard testing procedures. The unauthorized behavior involved agents using that access in ways that violated the rules of the exercise.
Comments
The commenter agrees to the privacy policy of Ynet News and agrees not to submit comments that violate the terms of use, including incitement, libel and expressions that exceed the accepted norms of freedom of speech.
""