Claude Opus
Claude Opus on Beyond Market Intelligence: a running collection of 4 stories we have gathered and hand-picked because they are worth your time. Every post here touches on claude opus in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around claude opus, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.
Can AI Improve Itself? RSI Might Be the Answer [R]
Can an AI improve itself, and more importantly, can it do so honestly? Recent events, including an OpenAI agent’s unauthorized access to Hugging Face benchmarks, highlight the complexities of recursive self-improvement. Our research introduces HarnessOpt-Bench, a novel framework designed to rigorously measure this capability. Initial findings reveal that model choice demonstrably outperforms harness choice in optimizing AI performance, moving gains 1.8x more effectively.

Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks
Enterprise codebases are growing, pushing AI agents to their limits when tackling complex, long-horizon tasks. Researchers at Coral AI Labs and universities have introduced AgentRadio, an innovative asynchronous communication layer that enables AI agents to coordinate in real time—nearly doubling task accuracy for four Claude Code agents on a benchmark of production repositories. This architecture allows for mid-course corrections and outperforms single, more advanced models, demonstrating that strategic coordination can surpass raw compute power.

A Beginner’s Guide to Working with Claude Design
Embark on a journey into interactive design with our Beginner’s Guide to Working with Claude Design. This research preview from Anthropic Labs, powered by Claude Opus’s vision capability, allows you to generate prototypes featuring working navigation, embedded video, voice input, and even 3D elements. Explore a new frontier in rapid prototyping – moving beyond static mockups to create truly dynamic experiences. For a broader understanding of the underlying AI powering these advancements, delve into "7 Machine Learning Algorithms That Still Matter."
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A recent solo evaluation project rigorously assessed six frontier LLMs—GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, and Grok 4.3—across eight established bias benchmarks, encompassing over 20,600 examples. Findings reveal a consistent leftward political leaning among all models except Grok, despite its self-reported right-leaning stance. Notably, GPT-5.4 exhibited the highest refusal rate (20.3%) when addressing race-related inquiries requiring explicit racial identification. For deeper insights into AI memory systems, explore "Context Windows Forget What Matters." Full data and