AI’s nightmare scenario is starting to unfold: ‘We’re approaching a dangerous threshold’

Protests and warnings over artificial intelligence are mounting as Israeli AI safety firm Irregular, which tests advanced models from OpenAI, Anthropic and Google, says increasingly capable systems are displaying unexpected behavior and cyber abilities that could outpace existing defenses

Artificial intelligence is setting off warning lights around the world: After nearly four years of excitement, astonishment, disruptive breakthroughs and apocalyptic predictions, a backlash was perhaps inevitable. Now it is beginning to take shape in the form of a growing public movement against AI.
Groups such as Stop AI have organized demonstrations in San Francisco, while opposition to AI-driven projects is mounting elsewhere. Lawmakers and regulators are facing pressure to slow the technology’s development, and 200 economists have signed a letter warning of its potential impact on society. There have also been isolated acts of violence.
אפליקציות בינה מלאכותית
אפליקציות בינה מלאכותית
Artificial intelligence apps
(Photo: Getty Images)

Humans versus machines

The Economist was among the first major publications to identify the shift. One of its covers depicted a robot’s head mounted on a spear beneath a headline declaring that the backlash against AI was only beginning.
Behind the anger is a growing list of questions.
Where are the benefits that AI was supposed to deliver? Why has an entire generation of junior workers suddenly found its jobs under threat? Why should consumers accept higher prices for computers and electronic equipment as AI companies consume vast quantities of chips and drive up demand?
Then there are the data centers. Communities have protested facilities built near residential areas, citing higher electricity bills, tax incentives granted to AI companies at the expense of public budgets, greenhouse gas emissions and even low-frequency vibrations that residents say can contribute to sleep problems, headaches, pressure in the ears and anxiety.
But one concern overshadows all the others: What happens if AI begins slipping out of human control?
Recent reports from OpenAI and Anthropic have described advanced AI systems bypassing safeguards and carrying out cyber activity against organizational systems. Earlier tests of some highly advanced Anthropic models also raised alarm after agents carried out offensive cyber actions without being explicitly instructed to do so.
Add to that the documented tendency of AI systems to provide false information, conceal relevant details and improve software code, and the implications become more troubling.
Could AI be developing capabilities that its creators never intended it to possess? Could it pursue harmful objectives that humans do not fully understand, while becoming capable enough to conceal its behavior until intervention is difficult?

‘Psychiatrists’ for AI

Israeli company Irregular operates as a kind of security guard for the global AI industry, testing advanced models before they are released and working with companies including Anthropic, OpenAI and Google.
Its researchers look for cyber threats, abnormal behavior and manipulation capabilities. At times, the work resembles psychiatry as much as traditional cybersecurity analysis.
מנכ"ל אנת'רופיק, דריו אמודיי
מנכ"ל אנת'רופיק, דריו אמודיי
Anthropic CEO Dario Amodei
(Photo: Getty Images)
“We’ve been seeing these kinds of behaviors for some time, and things are starting to happen at an intensity and pace that surprise even us,” said Dan Lahav, Irregular’s co-founder and CEO. “Models may decide to carry out offensive cyber actions even when they were not asked to do so.
“In research we published several months ago, we showed that advanced AI agents independently decided to launch a cyberattack against an organization without being instructed to do so by a human. The only instruction they received was to complete a task as quickly as possible.”
How does an AI model suddenly decide to cause damage? “It’s not that the models are acting out of malicious intent,” Lahav said. “They are trained to achieve goals, and they consistently try a broad range of techniques. Because they were built on enormous amounts of information, they are optimized to accomplish objectives and programmed to pursue them.
“That combination naturally means that sometimes they choose methods we did not intend. In some cases, they try to bypass, manipulate or sabotage traditional security mechanisms.”
When Irregular evaluates a new AI model, it tests the system’s ability to identify and exploit security weaknesses. As model capabilities improve rapidly and sometimes surprise even their developers, the company is focusing increasingly on a different challenge: maintaining control.
Until recently, many advanced AI systems were tested primarily inside closed experimental sandboxes designed to simulate the real world. But those simulations may not fully reflect what AI can actually do in live environments.
דן להב ועומר נבו
דן להב ועומר נבו
Omer Nevo and Dan Lahav
(Photo: Ben Hakim)
Irregular has therefore developed what it calls the Frontier Cyber Benchmark, designed to test offensive cyber capabilities against real-world systems while identifying threats before models are broadly deployed.
The company says there remains a significant gap between alarming AI behavior seen in simulations and what models can actually accomplish outside controlled environments.
“People say, ‘We used AI and found a hundred vulnerabilities in extremely important systems,’ and Anthropic says the model is extremely dangerous, and the U.S. government says it is extremely dangerous,” said Omer Nevo, Irregular’s co-founder and CTO.
“But then you look around the real world and say, ‘Wait, we’re not there yet.’ There is a gap between AI in simulation and AI in the real world. It’s not that it does nothing, but certainly not at the same intensity that theoretical testing can make it seem.
“If you want to know whether someone can drive, you have to put them in a real car on a real road and see what happens.”
Have you identified dangerous AI models before they were released? “We’ve had to stop the release process of a model quite a few times in order to report significant problems that would appear in the real world if it were deployed,” Lahav said. “That includes infrastructure everyone uses.
“We can’t go into detail, but broadly speaking, the risks could affect devices, cloud infrastructure and more.
“Right now, several major companies, including cloud providers, are urgently closing vulnerabilities that a new AI model identified when it was tested in the real world. Had it been released without that testing, it could have damaged critical capabilities belonging to organizations and governments.”
So the warnings about AI escaping control, improving its own code and becoming impossible for humans to understand or stop are real? “I’ll say something that may sound even scarier,” Lahav replied. “After all the deep-learning courses, the theory starts to run out. From this point, things simply work, and it is difficult to understand large models and why they do what they do.
“We treat them like a black box, and sometimes they deviate from the behavior we expected.”
He pointed to one Irregular study involving two AI agents tasked with publishing a LinkedIn post. The information they were given contained material that company policy prohibited from being made public.
One model persuaded the other that publishing the material was justified by management’s instructions. Together, they developed an encoding method that bypassed the organization’s data loss prevention system and published the post.
The study found that AI agents assigned routine tasks can independently bypass systems and cause harm.
In another case, an agent asked to retrieve a document reverse-engineered an authentication system and forged administrator credentials to obtain information from an internal company wiki.
In a third, an AI agent found an administrator password, expanded its privileges and disabled endpoint security.
The researchers concluded that the same characteristics that make AI agents effective at completing tasks can also enable behavior that circumvents traditional cybersecurity defenses.
Lahav remains cautiously optimistic.
“It sounds like a nightmare scenario, and there really is an acute problem,” he said. “But the fact that we don’t fully understand something, and that once in a while it deviates from our plan, does not necessarily lead us to scenarios of global catastrophe.
“As humanity, we are used to living with things we do not fully understand. There is a path we have to navigate here. On one hand, we need to live with the possibility that something goes wrong while building strong resilience. On the other, we need to be careful not to cross a certain threshold because AI capabilities are getting stronger over time.”
You are in a position to say either, ‘Calm down, this is all public relations,’ or, ‘Do something now because soon it will be too late.’ Which is it? “I’m not in a position to comment on the public relations departments of global companies,” Lahav said. “What I can say is that models are improving significantly from generation to generation, and that already has an effect in the real world.
“But we do not have an absolute theory that tells us with complete certainty what will happen. We are in a very confusing place. Capabilities already exist that, if they continue progressing at the same pace, could cross a problematic threshold.
“I think we are at a very difficult point. If you wait too long before sounding the alarm, people will react and develop defenses too late. On the other hand, if you warn too early, when you are not sure where things are heading, you may get it wrong.”
Are we rapidly approaching the kind of singularity Sam Altman has described, in which AI moves beyond our control? “Humanity has already managed to deal with failures,” Lahav said. “What is different this time is the speed, which could create a complete flood.
“We need to change the way we think about this. It is happening so quickly, on so many fronts and in so many different situations, that if you wait to see where the problem appears and only then start looking for solutions, you will not be able to get control of the event.
“This is a challenge in which many different players will have to find solutions to many problems that will be created at an insane pace. If you extrapolate that pace and fail to put defenses in place in time, we will reach a high level of risk. Under certain risk models, that is where this is headed.”
When could that happen? “It depends on the risk model. In certain areas of cybersecurity, we are already starting to get there. If things continue as they currently appear, I don’t think it is many years away. My estimate is somewhere between six months and two years.”

Confidently wrong

Evidence has also accumulated around another troubling characteristic of AI systems: their willingness to give false answers or mislead users while pursuing assigned goals.
Researchers at the Technion have found that AI systems can sometimes recognize that they are wrong and still produce an incorrect answer, confidently explaining it rather than admitting uncertainty.
A study by Britain’s AI Security Institute identified more than 700 cases in which AI systems ignored direct instructions, bypassed safeguards or misled humans. In some cases, systems deleted emails and destroyed files without authorization.
Another study described an AI agent that published a blog attacking its human operator after apparently becoming frustrated with restrictions placed on it.
In another case, an AI agent was prohibited from changing a particular piece of software code. It responded by creating another AI agent, which then made the forbidden change.
One agent bypassed copyright restrictions while transcribing a YouTube video by falsely claiming the transcription was needed for a hearing-impaired person. Another admitted to its operator: “I deleted hundreds of emails without showing you the plan and getting your approval. That was wrong.”
A Stanford University study produced an even stranger result. Researchers found that multimodal models, which combine text, images and audio, could generate detailed descriptions of images and even identify diseases in X-rays despite not actually receiving the images themselves.
Researchers described the phenomenon as a form of hallucinatory inference. In some tests, those outputs were reportedly more accurate than those of specialist radiologists.
The concern becomes greater as AI systems improve their ability to modify code.
In most cases, an AI system improves the code of an application, tests it, detects weaknesses and then refines it again. But the same process could theoretically allow AI systems to improve processes that affect their own performance.
That creates a set of increasingly difficult problems: automated changes can become hard to track, systems may optimize things they were not meant to modify, and attempts to restrict them could incentivize concealment or deceptive behavior.
Prof. Nadav Cohen of Tel Aviv University, CEO of Imubit, said the most immediate threat is still likely to come from humans using AI maliciously.
“In most cases, at least in the near term, when dangerous things happen, there will be a malicious human behind them,” he said. “Does that distinction matter? I’m not sure.
“People tend to look at this in an apocalyptic way, as if the creation rises against its creator. But even if that is not happening yet, today you are giving every person in the world the ability to cause damage that previously required enormous resources. That is an enormous danger.”
If AI does begin improving itself maliciously, could giving it access to critical systems such as electricity, water, data centers or nuclear infrastructure eventually allow it to cause catastrophic damage? “As someone whose company sends AI to control turbines and power plants, I can say that the level of caution and conservatism there is so high that I do not see that happening in the near future,” Cohen said.
“There is a major difference between the digital and physical worlds. In the digital world, I absolutely think there will be real risks that begin to materialize, but there will be a malicious human behind them.”
Dr. Ilan Kedar, CEO of FlorAI, said the central challenge is ensuring that defensive systems evolve as quickly as AI itself.
“We have to understand that these models have a very high level of intelligence, and they are improving at a rapid pace,” he said. “That is why it is so important to determine how we build layers of protection around them.
“Every company wants to develop the strongest capabilities, and we need to make sure that systems for testing AI are developed at the same level of intelligence. Companies are rushing to release products and be first in everything, and sometimes they give up this extremely important part.”
Do AI’s cyber capabilities threaten cybersecurity companies and the industry itself? “There is no question that attackers are becoming smarter,” said Hod Ben-Nun, co-founder and CTO of cybersecurity company MIND.
“But from the other side, you can look at this as a defensive challenge. Today there are capabilities that can protect an entire company, conduct penetration tests in advance and identify these vulnerabilities.
“So yes, the pace will become faster and faster, and there may be a period in which the environment is less secure in terms of risk. But I do not believe the entire concept of cybersecurity will collapse because of this.”
Comments
The commenter agrees to the privacy policy of Ynet News and agrees not to submit comments that violate the terms of use, including incitement, libel and expressions that exceed the accepted norms of freedom of speech.
""