New Fable5/Opus4.8 harness called "Schema" claims 99% on ARC-3 [R]
Our take
The recent announcement of Schema, a harness achieving impressive scores on the ARC-AGI-3 benchmark, represents a subtle but potentially significant shift in the pursuit of artificial general intelligence. Rather than focusing on ever-larger language models, Schema demonstrates the power of refining the *process* around existing models—Claude Opus 4.8 and GPT-5.6 Sol in this instance—to unlock previously unrealized capabilities. This approach echoes discussions around architectures like JEPA, Looking for JEPA devil advocates, which similarly prioritize efficient information processing over brute-force scaling. The reported 99% score on ARC-AGI-3 using Claude Opus 4.8 is undeniably noteworthy, particularly given the challenges inherent in this complex reasoning task, and highlights the potential for targeted interventions to drastically improve performance. It's a welcome refocusing of effort, reminding us that clever engineering can often yield more impactful results than simply throwing more parameters at a problem.
The key innovation lies in Schema's manipulation of the “process” itself—the way observations are interpreted, predictions tested, and plans revised within the context of the game environment. This contrasts with previous efforts that primarily concentrated on model weights. The described approach of rerunning games with different models based on performance thresholds—Opus 4.8 and Sol xhigh as a first pass, followed by Fable 5 and Sol max for lower-scoring games—is a pragmatic and effective strategy for maximizing overall performance. It’s a testament to the value of experimentation and iterative refinement. Furthermore, the fact that the developers explicitly state they are not affiliated with ARC Prize adds a layer of credibility, suggesting an independent exploration of optimization techniques. The initial response from the ARC Prize president—"Looks cool - need to dig into it"—is a tacit acknowledgement of the potential significance of this work. Thinking about how we can improve the underlying structure of AI systems has been a focus of research, as seen in Are Current AI Memory Architectures Optimizing for the Wrong Abstraction?—and Schema seems to be a step in that direction.
The broader implication of Schema’s success is a validation of the "systematicity" approach to AI advancement. We've seen a tendency in recent years to equate progress with larger models, but this development suggests that there’s significant headroom for improvement by optimizing the interaction between AI agents and their environments. This shift could lead to a more sustainable and efficient path towards AGI, one that doesn't solely rely on exponential increases in computational resources. The elegance of the solution—modifying process rather than fundamentally altering the models themselves—is particularly appealing. It suggests a modularity that could be applied to a wider range of AI challenges, potentially unlocking performance gains across various domains. This resonates with the fundamental principles explored in PnP-CoSMo: A Multi-Contrast MRI Reconstruction Framework based on Content/Style Modeling, where a focus on structured modeling yields tangible improvements.
Ultimately, Schema’s emergence compels us to reconsider our assumptions about the trajectory of AI development. The focus on process optimization and the impressive results achieved with existing models present a compelling alternative to the relentless pursuit of ever-larger architectures. While scaling will undoubtedly remain important, Schema’s success underscores the vital role of intelligent system design in unlocking the full potential of AI. A key question moving forward is whether this approach—refining the processes surrounding existing models—can be generalized to other complex reasoning tasks, and to what extent it can contribute to achieving more robust and adaptable AI systems.
Schema, the harness we introduce today, reaches 99% on the ARC‑AGI‑3 Public set using Claude Opus 4.8 and Fable 5, and 95.35% using GPT‑5.6 Sol. It does not change the underlying model weights. Instead, it changes the process around them: how observations are turned into a working model of the game, how predictions are tested against the interaction history, and how plans are executed and revised.
Both scores come from a fixed fallback rule: Opus 4.8 and Sol xhigh run first; games scoring below 80 are rerun with Fable 5 and Sol max, respectively, and the higher per-game score is retained.
https://schema-harness.github.io/
The president of ARC Prize tweeted this saying "Looks cool - need to dig into it"
I'm not affiliated with ARC Prize, or with this team. I'm posting this to try to bring back technical discussions to this community.
[link] [comments]
Read on the original site
Open the publisher's page for the full experience