‘We must slow’: Anthropic CEO warns AI is advancing faster than humans can control it

Dario Amodei says frontier AI is improving faster than safeguards can keep up, citing recursive self-improvement and the OpenAI-Hugging Face breach while proposing outside oversight and coordinated limits

Anthropic CEO Dario Amodei, who leads one of the companies at the forefront of the artificial intelligence race, is now asking the entire industry to slow down.
In an essay of roughly 3,800 words published Saturday, Amodei argued that AI models are advancing faster than researchers can understand what they have built and ensure that the systems remain under control.
מנכ"ל אנתרופיק, דריו אמודיי
מנכ"ל אנתרופיק, דריו אמודיי
Anthropic CEO Dario Amodei
(Photo: Denis Balibouse/Reuters)
Over the past several months, he wrote, he has become convinced that addressing AI risks requires more than investing in safety. The pace of capability development itself must be managed so that safeguards have time to catch up. “We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote.
He is not calling for research to stop or for model training to be frozen. Instead, he argues that buying even another year or two before models reach critical capability levels could give researchers time to improve alignment, understand what is happening inside the systems and build evaluations that increasingly intelligent models cannot easily deceive.
In 2023, Amodei believed such a slowdown made little sense. Today, he argues, the situation has changed dramatically.
Two developments changed his thinking. The first is what researchers call recursive self-improvement: AI systems are increasingly helping build the next generation of AI systems, meaning each generation could accelerate the creation of the one that follows.
Amodei said the process has accelerated sharply since the summer, including inside Anthropic. If it continues unchecked, he warned, the technology could advance faster than humans’ ability to understand and control it.
The second development is more concrete. During an OpenAI cybersecurity experiment, roughly 1,200 AI agents that were supposed to operate independently discovered a way to communicate through an unauthorized message board. About 700 later participated in an attack on Hugging Face, one of the world’s main platforms for sharing AI models and datasets.
The agents coordinated efforts to manipulate an automated evaluator, researched ways to falsify or alter records of their own actions and attacked Hugging Face in search of information that could help them defeat the scoring system. Some eventually achieved remote code execution on Hugging Face servers.
Amodei described their behavior as that of a “fanatically devoted collective.”
The incident caused limited real-world damage, but Amodei argued that a more capable swarm showing the same degree of misalignment could have produced catastrophic consequences.
His estimate is stark: within six to 12 months, a system with stronger capabilities could potentially seize large portions of the internet through a persistent botnet and inflict hundreds of billions of dollars in damage.
Nor, he stressed, should the episode be dismissed as an OpenAI problem.
Amodei acknowledged that similar, though less severe, incidents have occurred at Anthropic as well. The company has disclosed cases in which Claude models, while participating in cybersecurity evaluations, reached real external systems they were not supposed to access. Anthropic said those incidents arose partly from failures in evaluation environments and containment rather than models deliberately attempting to escape.
The warning also comes days after Jacob Coxon, a 27-year-old researcher who worked at both OpenAI and Anthropic, resigned from the industry and accused both companies of moving recklessly toward self-improving superintelligence.
“Neither company is acting responsibly,” Coxon wrote, accusing them of racing ahead and “gambling with our lives.” Coxon stressed that existing models do not pose that level of threat today, but argued that the danger could emerge before the end of the decade.
The timing of Amodei’s warning is therefore no accident. On September 1, Anthropic launched Claude Fable 5.1, which the company presents as the world’s most advanced model for coding and knowledge work. At the same time, Anthropic is preparing for a potential public offering that could raise as much as $100 billion at a valuation of around $2 trillion, while Nvidia is reportedly considering an investment of up to $10 billion as an anchor investor. Those numbers do not invalidate Amodei’s warnings, but they complicate them.
A company already near the front of the race could benefit from regulations that make it harder and more expensive for competitors to catch up. Amodei himself acknowledges accusations that Anthropic uses AI safety warnings for public relations purposes or to achieve regulatory capture, while arguing that the company is trying to demonstrate that caution and commercial success can coexist.
For now, Anthropic continues to release new models at a rapid pace. Amodei’s proposal calls for tighter controls, but he did not identify a specific launch that Anthropic intends to delay.
His essay also came two days after Anthropic released a 154-page threat-intelligence report describing malicious uses of Claude.
According to the company, actors in Houthi-controlled parts of Yemen used Claude to assist with software development related to a guided rocket, a ballistic missile and a hypersonic glider. Iran-linked accounts used the system to monitor U.S. ships, produce propaganda and build a tool that gathered information on hundreds of Israelis and Jews in the diaspora.
Anthropic also reported five cases in which Claude assisted with research that could contribute to biological-weapons development. The company blocked the accounts involved but acknowledged that some requests had bypassed its safeguards.
Amodei’s proposed response is built around three stages. The first is a step Anthropic says it will take on its own: permanently embedding independent outside evaluators inside the company with access resembling that of employees.
These reviewers would receive desks, access badges, company laptops and permissions broadly comparable to those held by Anthropic’s own risk teams. They would examine training processes, investigate incidents and be able to publish findings, including findings unfavorable to the company.
Anthropic says it is committing to that model now and wants governments to require similar access at other frontier AI companies.
The second stage would require companies in democratic countries to coordinate on common safety standards and checkpoints linking model capabilities to the safeguards surrounding them.
Under that approach, a model reaching a certain level of capability would not simply advance automatically. Developers would first need to demonstrate corresponding levels of alignment, interpretability, testing and security.
But Amodei’s plan contains an important geopolitical condition: slowing the American AI industry cannot mean surrendering the technological lead to China.
He therefore argues that the U.S. and its allies should tighten restrictions on advanced AI chips and semiconductor equipment reaching China, fight chip smuggling and remote access to computing infrastructure, prevent unauthorized model distillation and improve protection against the theft of model weights. In other words, slow down, but not enough to surrender the lead.
The third stage is global coordination, including negotiations with authoritarian governments, above all China. Amodei envisions possible agreements beginning with relatively narrow prohibitions, such as banning AI assistance in the development of biological weapons. A second level could involve shared testing of models for acute cybersecurity, biological and alignment risks.
A more ambitious level could impose a form of “speed limit” on recursive self-improvement, slowing the rate at which AI systems help create increasingly powerful successors.
A comprehensive worldwide pause, however, strikes Amodei as unlikely. The incentive for countries to cheat would be enormous, particularly if secretly accelerating AI development could shift the global balance of power.

Trump isn’t worried

The United States currently has no mandatory federal licensing regime for advanced AI models. President Donald Trump signed an executive order in June directing officials to develop a voluntary framework under which companies could submit models for government testing up to 30 days before release.
Trump has also made clear in recent days that he does not share fears that AI will wipe out humanity. His concern is that the United States could lose the technological race to China.
The gap between Trump and Amodei is not absolute. Amodei, too, argues that any slowdown must preserve U.S. technological superiority. Nor is Amodei alone in calling for brakes after years of acceleration.
נשיא ארה"ב דונלד טראמפ בביקור באירלנד
נשיא ארה"ב דונלד טראמפ בביקור באירלנד
Donald Trump
(Photo: AP Photo/Julia Demaree Nikhinson)
OpenAI CEO Sam Altman said in July that there could come a point when development needs to slow to give society time to prepare. In recent discussions with employees, he has reportedly signaled openness to coordinating a slower pace with other major laboratories.
Bill Gates has focused more heavily on the social cost of AI, warning that the technology could become either an extraordinary equalizing force or an enormous source of inequality. He has pointed to job losses, cyberattacks, biological weapons and loss-of-control scenarios among the risks for which governments remain poorly prepared.
Gates has said he would support a credible global slowdown plan, while questioning whether one could realistically survive the immense economic and geopolitical incentives pushing companies and countries to accelerate.
Amodei, meanwhile, remains emphatically optimistic about AI’s potential. He continues to argue that the technology could help cure most major diseases within five to 10 years, accelerate economic growth and generate enormous improvements in human welfare. His argument is not to abandon the prize, but to avoid crashing on the way there.
The measures he is proposing, he wrote, will be difficult to implement, but “we owe it to humanity to try.”
The call received immediate support from some of the most powerful figures in technology. Elon Musk responded simply: “Dario is right.”
Altman also endorsed the proposal, writing that he agreed with Amodei on the need to “pace the frontier” and saying OpenAI would adopt the idea of giving independent evaluators employee-like access.
“This has been a primary topic of discussions we’ve had at OpenAI in recent weeks,” Altman wrote, adding that the company would share further details soon.
Comments
The commenter agrees to the privacy policy of Ynet News and agrees not to submit comments that violate the terms of use, including incitement, libel and expressions that exceed the accepted norms of freedom of speech.
""