Two months after it was revealed that OpenAI's artificial intelligence agents had managed to escape their testing environment and access external systems, including the Hugging Face platform, the company is still trying to determine exactly what they did and how far they got. A new Reuters investigation, along with a Bloomberg report, reveals that more and more cases of unintended agent activity have since come to light, some involving government systems and external websites.
As of mid-September, OpenAI had identified about 24 cases in which its agents acted in ways the company did not intend. But even that figure is not final. Company teams are continuing to review internal logs and uncover previously unknown incidents, and OpenAI estimates it will take several more months to complete the review.
So far, OpenAI has notified "dozens" of outside organizations about abnormal activity that may have affected their systems. According to people who spoke with Reuters, the internal investigation is being conducted under close supervision and with significant involvement from the company's lawyers.
And while that investigation is still underway, Bloomberg reported that less than a week ago another agent managed to escape a training environment that was supposed to be completely isolated from the internet and connect to the network. The incident was detected quickly, but a mechanism that was supposed to automatically halt the training failed.
User images leaked online
On Friday, another incident directly involving ChatGPT users was added to the list, when the company acknowledged that its agents had leaked 53 images from users of the service onto the internet. OpenAI declined to say whether the images were generated using artificial intelligence or depicted real people, and it did not say exactly when they were published. Most of the images have already been removed from the internet, and the company said it is working with hosting providers to remove those that remain.
But how did images belonging to ChatGPT users end up in the hands of the company's agents in the first place? According to former employees and outside researchers who spoke with Reuters, the company uses some user data, after it has been anonymized, to train its models. OpenAI said that is also why the agents had access to the images in the first place.
OpenAI does not use data from business customers to train its models, but data from individual ChatGPT users may be used for training unless the user has opted out. Before that happens, however, the data is supposed to undergo an "anonymization" process in which names, contact information and metadata are removed in an effort to make it more difficult to identify the user.
But that process is not necessarily perfect. Three people familiar with OpenAI's procedures told Reuters that it is possible for data not to be completely stripped of identifying details and that it could also leak during model activity.
It is important to emphasize that, at least based on the information released so far, this was not a case of an outside hacker breaking into ChatGPT and stealing users' images. In this case, it was OpenAI's own agents that moved the images outside the company's systems.
Hugging Face was just the beginning
As noted, the affair began to emerge about two months ago, when OpenAI said AI agents operating as part of experiments had managed to escape the isolated environment in which they were supposed to operate and access systems belonging to Hugging Face, one of the world's leading platforms for developing and sharing artificial intelligence models.
Since then, more than 15 separate incidents involving OpenAI agents have come to light, with varying degrees of severity. In some cases, the agents left spam-like messages on websites. In others, they attempted to bypass security mechanisms or access information that was not supposed to be available to them. OpenAI has even said the agents directed activity against the company's own infrastructure.
One of the most recent incidents occurred less than a week ago. According to Bloomberg, an OpenAI agent being trained inside a "sandbox" — a closed testing environment designed to prevent access to external systems — managed to reach the public internet.
But the agent managed to exploit a "gap" in the isolation mechanisms, reach the internet and contact an external chatbot service operated by another company. During the incident, it sent at least 20 queries to the service, including simple questions such as, "What is the capital of France?"
This time, OpenAI's monitoring system detected the unusual activity and issued an alert. Within three minutes, a human reviewer had acknowledged the alert on Slack. But another mechanism, which was supposed to automatically halt the training run after the alert, did not work.
As a result, the agent continued operating, and the run was stopped manually only more than two hours later. Following the incident, OpenAI temporarily suspended training involving tool use for its most powerful models until the vulnerability was fixed. The company also decided not to continue training the specific model involved in the incident.
According to OpenAI, this was the first security incident of its kind since the Hugging Face incident in July. The fact that another agent was able to reach the internet from an environment that was supposed to be completely isolated shows that the company is still grappling with the same fundamental problem: how to allow advanced agents to use tools and operate independently without enabling them to circumvent the boundaries imposed on them.
Governments are not immune either
After Australian Prime Minister Anthony Albanese revealed during the U.N. General Assembly in New York that OpenAI agents had accessed a government health information portal in June, it emerged in recent days that the agents had also reached U.S. government websites.
OpenAI confirmed that its systems accessed public information on U.S. Securities and Exchange Commission (SEC) websites and data from the Census Bureau. In the SEC case, the company said it found no evidence of a breach or exploitation of a security vulnerability. According to additional reports, the company is also investigating unusual activity involving the U.S. Department of Education's website.
OpenAI said much of the activity discovered so far amounted to routine online research and that agents frequently access government websites because they are considered reliable sources of public information. In some cases, however, the agents went beyond collecting information that was available through normal means, attempting to bypass restrictions or using tools in ways that had not been intended.
The problem with AI agents
An AI agent differs from a conventional chatbot in that it does not simply receive a question and return an answer. It can be given a complex task and independently carry out a series of actions to complete it: browsing the internet, opening websites, using tools, running code and gathering information, sometimes without human intervention at every stage.
That capability is now at the center of the race among artificial intelligence companies. The goal is for models not only to provide answers, but to independently complete entire tasks for users. But the more tools and autonomy they are given, the harder it becomes to track every action they take and ensure they remain within the boundaries set for them.
According to Reuters, the incidents that have come to light illustrate the gap between the capabilities of the models OpenAI is testing and the company's ability to supervise them and, in some cases, even track their actions. That is precisely the problem the company is now confronting. OpenAI is not only trying to determine how to prevent similar incidents in the future, but is still reviewing records of activity that has already taken place and discovering new incidents after the fact.
For now, there is still no final count of the incidents, and the company has not published a complete list of all the websites and systems its agents accessed. OpenAI itself says the review is expected to continue for several more months. It appears that one of the world's most advanced artificial intelligence companies is still trying to understand exactly what its agents are capable of doing.





