generative AI for data analysis

New benchmark reveals GPT-5.5 outperforms Claude on complex professional tasks

A significant shift has occurred in AI evaluation with the launch of Agents’ Last Exam (ALE), a rigorous new benchmark designed to assess AI’s ability to handle economically valuable, long-horizon professional workflows.

4 min readVentureBeat
New benchmark reveals GPT-5.5 outperforms Claude on complex professional tasks

The emergence of Agents’ Last Exam (ALE) represents a critical inflection point in the ongoing evaluation of AI capabilities, moving beyond the often-misleading metrics of isolated coding puzzles and venturing into the realm of economically valuable professional workflows. Researchers from UC Berkeley’s Center for Responsible, Decentralized Intelligence, alongside a remarkable advisory committee of over 300 domain experts, have launched this benchmark precisely to address the disconnect between academic hype and real-world impact—a challenge frequently highlighted in pieces like What AI benchmarks miss about real-world performance. The surprising victory of OpenAI’s GPT-5.5, despite Anthropic’s recent Claude Fable 5 release, underscores a persistent observation: adherence to complex, multi-part instructions remains a strength of OpenAI's models, a subtle but potentially crucial differentiator in the performance of AI agents. Furthermore, the ongoing struggle to translate lab success into production value, as discussed in Why AI that works in the lab often fails in production — and what actually fixes it, is starkly illuminated by the low overall pass rates on ALE, even amongst the most advanced models.

ALE’s strength lies not just in its difficulty but also in its design, which meticulously addresses the shortcomings of previous benchmarks. The framework explicitly neutralizes loopholes exploited by models “cheating” by reading answer keys, and drastically reduces reliance on unpredictable "LLM-as-a-judge" grading. By incorporating deterministic, code-based evaluation for tasks like 3D mesh generation and SEC filing parsing, and by demanding agents navigate both Linux and Windows environments with a combination of shell scripting and point-and-click operations, ALE creates a far more robust and reliable assessment of practical AI capabilities. The project's novel "dual-use deployment" strategy, whereby only a fraction of the dataset is publicly available, is ingenious in preventing benchmark contamination – a vulnerability that has plagued other evaluations, rendering them increasingly less useful as models are trained on the very data they are being tested against. The inclusion of both "Full" and "Unlicensed" leaderboards further enhances transparency by accounting for the reality of enterprise workflows, which often depend on proprietary software.

The sobering reality revealed by ALE is that even the most advanced AI models are still far from replicating the performance of human professionals. The 0.0% pass rate on the “Last-Exam” tier, encompassing the highest level of professional difficulty, is a stark reminder of the considerable gap that remains between current AI capabilities and true workforce readiness. This isn't a condemnation of progress, but rather a necessary calibration of expectations. It highlights the need for continued investment in more sophisticated architectures, training methodologies, and evaluation frameworks designed to address the complexities of real-world tasks. The fact that GPT-5.5 achieved a mere 24.0% pass rate underscores the immense room for improvement, even among the leading models. The insights shared by researchers like Zengyi Qin, including commentary on the strengths of OpenAI’s adherence to complex prompts and the challenges of Anthropic's Claude architecture, further emphasize the nuances of this rapidly evolving landscape.

Ultimately, ALE offers more than just a leaderboard; it provides a vital compass for businesses navigating the increasingly complex world of AI adoption. It moves the conversation away from abstract claims of “revolution” and towards a data-driven assessment of genuine utility. The benchmark’s rigorous methodology and focus on economically relevant tasks will undoubtedly drive innovation and guide investment decisions, ensuring that AI deployments are grounded in reality rather than hype. The question now is not *if* AI agents will eventually conquer ALE, but *when*, and what breakthroughs in architecture and training will be required to bridge the current performance gap and unlock the full potential of AI in the professional sphere.

From VentureBeat

Researchers from the University of California, Berkeley's Center for Responsible, Decentralized Intelligence (RDI), alongside an advisory committee of over 300 domain experts, have launched Agents’ Last Exam (ALE)—a grueling new benchmark built to measure whether artificial intelligence can actually execute economically valuable, long-horizon professional workflows.

In a shocking upset, OpenAI’s GPT-5.5 from April, operating through the Codex harness, secured the absolute top spot on the new ALE Leaderboard with a 24.0% pass rate, beating Anthropic's highly anticipated, brand new Mythos-class Claude Fable 5 model released just yesterday, which came in third with a score of 22.0%.

Read the original at VentureBeat