A new study by researchers at the University of California, Berkeley, found a wide gap between the promises made by artificial intelligence companies and the technology's performance on complex real-world workplace tasks.
The test, called Agent’s Last Exam and led by a team from Berkeley’s Center for Responsible, Decentralized Intelligence, examined the ability of the most advanced language models on the market to handle complex professional tasks in real-world settings.
Gallery


No model managed to cross the 25% success threshold across the full set of assigned tasks
(Photo: Getty Images)
Clear-cut results
The study’s findings painted a clear picture: No model managed to cross the 25% success threshold across the full set of assigned tasks. OpenAI’s ChatGPT-5.5 ranked first with a low success rate of just 24%, while other advanced models, including Anthropic’s Fable 5 and Cursor’s Composer 2.5, failed on most of the tasks they were given.
In the ranking of the most difficult tasks, which involved processes requiring continuous reasoning, broad knowledge and reliable execution over time, all of the tested models recorded a zero percent success rate.
The Berkeley test included more than 1,500 tasks designed by experts from 55 professional fields, including finance, law and traditional industries. The researchers emphasized that the systems’ main weakness does not stem from a lack of data, but from a lack of practical experience managing complex workflows that require flexibility, the ability to change strategy while working and adaptation to specific failures. In other words, these systems struggle to deal with uncertainty and changes that emerge after the fact.
The findings become even more significant when compared with the standard tests that have been used by the industry in recent years. While AI models have delivered impressive results, and in some cases surpassed human performance, on theoretical academic benchmarks such as MMLU or on narrowly focused basic programming problems, their performance collapses almost entirely when a task requires a long chain of decisions without precise instructions provided in advance.
Gap between tech companies and reality on the ground
An analysis of global trends shows a perception gap between reports coming from technology companies in China and the United States and the reality on the ground. While organizations in Europe have shown greater caution and adopted strict regulatory procedures in response to concerns about model reliability, American and Chinese technology giants continue to promote claims that the transition to fully autonomous AI agents is accelerating.
The current study reinforces an argument recently raised in regulatory discussions in Washington and Brussels: the lack of a uniform, independent evaluation standard allows companies to present partial data that may mislead the public and policymakers.
Modern AI agent technology is based on advances in deep neural networks first introduced in the previous decade. Similar to previous technological revolutions, such as speech recognition systems in the 1990s or early attempts to develop fully autonomous vehicles in the last decade, the initial breakthrough phase is characterized by a dramatic leap in basic capabilities that is quickly matched by inflated expectations.
Only in the second stage, when systems are required to match the narrow margins of error tolerated in human labor, do the technology’s structural limitations become apparent.
However, experts in the field estimate that the immediate losers from these changes will not be workers in complex professions requiring flexible decision-making abilities, but rather those engaged in routine, repetitive tasks. Structured processes that generate large amounts of accessible data will continue to be automated rapidly, while roles requiring multidisciplinary skills and critical thinking will remain relatively protected from full replacement by machines in the near future.


