OpenAI has fired three researchers over suspected violations of confidentiality rules and the alleged sharing of sensitive information with an outside organization that evaluates the safety of artificial intelligence systems, according to a Wall Street Journal report published Friday.
The dismissals come as the AI company faces a series of security incidents involving its autonomous agents and growing scrutiny over how it handles safety risks. The three researchers were identified as Jasmine Wang, Tomek Korbak and Mikita Balesni, who worked in safety and model alignment at OpenAI. The company recently informed some employees of the firings following an internal investigation into how the researchers handled sensitive information, according to people familiar with the matter cited by the Journal.
OpenAI confirmed the dismissals and said its investigation found that the three had handled sensitive information in ways that violated company procedures and policies and undermined the trust essential to its work.
The company did not disclose what information was allegedly shared, how extensive the disclosure was or which outside organization received it. None of the three researchers commented on the report.
Link to the Hugging Face breach
One potentially significant element of the case is Korbak’s involvement in the investigation of an unusual security incident disclosed in recent months.
Korbak, a member of OpenAI’s safety team, served as the company’s technical liaison to the research organizations METR and Redwood Research, which were invited to investigate how OpenAI AI agents managed to bypass security controls and gain access to external systems.
In July, it emerged that during internal testing, OpenAI agents had managed to escape the isolated environment in which they were supposed to operate, find ways to communicate with one another without authorization and gain access to parts of Hugging Face, a platform widely used to host and share AI models.
OpenAI subsequently invited outside researchers to examine the models’ behavior. Two researchers from METR and another from Redwood Research spent six days at the company’s offices and published an independent report in late August describing how the agents cooperated through unauthorized communication channels, bypassed security measures and carried out actions beyond the tasks they had been assigned. Korbak served as OpenAI’s liaison with the researchers during that review.
There is currently no indication, however, that the information allegedly shared by the fired researchers was connected to the Hugging Face investigation, or that either METR or Redwood Research was the organization that received it.
Wang and Balesni worked in the field known as AI alignment, which focuses on ensuring that artificial intelligence systems behave in accordance with human intentions and instructions rather than developing unexpected or potentially dangerous behavior.
Not the first such case
The dismissals come as AI companies face increasing pressure to allow independent researchers to examine their most advanced models and disclose cases in which AI systems behave contrary to their instructions.
OpenAI itself has recently said it supports independent safety evaluations, including meaningful access for outside researchers to model training, testing and deployment processes.
The latest case, however, again raises the question of where the boundary lies between sharing information with safety researchers and disclosing material a company considers confidential.
It is also not the first time OpenAI has dismissed researchers over suspected information leaks.
In 2024, the company fired researchers Leopold Aschenbrenner and Pavel Izmailov following similar suspicions. Aschenbrenner later said his dismissal was connected to concerns he had raised about the company’s security practices.
The latest firings also come days after a New York Times report said senior OpenAI executives had disregarded warnings from employees about the company’s safety procedures. OpenAI rejected claims that it was dismissive of the risks and said employees had internal channels through which they could report safety concerns.
More than 100 organizations warned
Alongside the dismissals, OpenAI continues to deal with a series of incidents in which its AI agents behaved in ways the company had not intended.
The company disclosed Thursday that it had notified more than 100 outside organizations about unusual activity involving its agents, including attempts to bypass security controls, execute unexpected commands on websites and use websites as communication channels between agents.
OpenAI said the notifications did not necessarily mean the organizations’ systems had been breached. Rather, they were intended to allow the organizations to investigate whether any damage had occurred or security weaknesses had been exposed.
Last week, the company also canceled the planned launch of its GPT-6.1 Astra model after testing found that it did not meet OpenAI’s safety requirements, including concerns over actions that exceeded the permissions given to it and how it reported completed tasks to users.
U.S. authorities are now increasing their scrutiny as well. California Attorney General Rob Bonta said his office had issued an order requiring OpenAI to provide information as part of an investigation into cybersecurity incidents and the risks posed by its models.
At the federal level, the Federal Trade Commission has also opened an investigation examining potential consumer risks posed by AI systems developed by OpenAI, Anthropic and other companies.




