benchmarks

benchmarks on Beyond Market Intelligence: a running collection of 13 stories we have gathered and hand-picked because they are worth your time. Every post here touches on benchmarks in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around benchmarks, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet
VentureBeat

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

Meta's newest AI model, Muse Spark 1.3, delivers notable performance gains over its predecessor, achieving "frontier performance" as CEO Mark Zuckerberg proclaimed. While the most impressive results stem from a "max reasoning" configuration still undergoing safety testing, the broadly available version ranks among the strongest price-performance offerings near the top of independent model evaluations. Though not currently leading the leaderboard—Anthropic’s Claude Fable 5.1 still holds that distinction—Muse Spark 1.3 represents a significant step forward, trading wins with OpenAI and Anthropic on key coding benchmarks.

Machine Learning

Sliding-window attention beats linear on long-context reasoning [R]

Recent research challenges the prevailing trend of post-training linear attention models in large language models. A new preprint demonstrates that Sliding Window Attention (SWA), a simpler and computationally efficient fix for the quadratic cost problem, consistently outperforms linear variants—often by a factor of 2 to 10 on long-context reasoning benchmarks like Needle-in-a-Haystack and BABILong. The authors assert that SWA represents a superior baseline, requiring no post-training and offering significant memory advantages.

Machine Learning

Where to submit stat/prob ML [D]

The dominance of large language models (LLMs) at top machine learning conferences has prompted a critical question: where does the statistical and probabilistic machine learning community find its home? While venues like NeurIPS and ICLR now largely focus on agentic LLM applications, researchers like Arnaud Doucet, Aapo Hyvärinen, and others continue to publish impactful work. AISTATS and UAI appear increasingly viable options, offering a more focused platform for stat/prob ML advancements.

An Anthropic researcher just gave us a peek at self-improving AI
TechCrunch

An Anthropic researcher just gave us a peek at self-improving AI

Recent advancements demonstrate the remarkable potential of self-improving AI. An Anthropic researcher recently showcased a system that successfully addressed ten distinct benchmarks for misaligned behaviors – achieving performance gains across all areas without compromising overall function. This signifies a crucial step toward safer and more reliable AI. Explore this progress and the broader landscape of AI development; for deeper insights into maximizing AI agent performance, see our article, "Connecting My LangGraph AI Agent to Postgres."

Bart- A vintage llm [R]
Machine Learning

Bart- A vintage llm [R]

Unbounded Labs proudly introduces Bart, a 2.82B parameter LLM meticulously trained from scratch on a unique corpus of 20.1B tokens of English text predating 1931. After three months and a modest $800 investment, we’ve achieved a significant milestone: the best-performing vintage base model at its scale on Vintage CORE. Our research, detailed in a comprehensive article, explores the potential for LLMs to replicate historical scientific reasoning—a crucial step toward understanding AI originality. Explore Bart and our methodology at the links provided.

Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model
TechCrunch

Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model

The AI community is buzzing: Z.ai, a previously enigmatic AI lab, has confirmed its development of Ox Alpha, the open AI model currently dominating benchmarks and leaderboards. This marks a significant development in accessible AI research, with the model's weights slated for imminent release. Z.ai’s emergence highlights the accelerating pace of innovation within the field. For a broader perspective on AI’s societal impact, explore our article "Understanding the Impact of AI on Job Markets" for insights into how AI is reshaping the future of work.

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
TechCrunch

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI’s new Jalapeño chip represents a significant advancement in AI inference capabilities. Benchmarks from SemiAnalysis’ InferenceX demonstrate Jalapeño’s exceptional performance, registering both more tokens per user and superior throughput per kilowatt compared to current state-of-the-art solutions. This positions Jalapeño as a leader for fast, scalable AI deployments. Explore the broader landscape of AI memory and its implications—similar to Anthropic’s recent enhancements to Claude, as detailed in "Claude Cowork finally remembers what you told the app in chat."

I built an open-source roguelike specifically for training game-playing agents [P]
Machine Learning

I built an open-source roguelike specifically for training game-playing agents [P]

For researchers and AI practitioners seeking a streamlined environment for reinforcement learning agent training, meet DelveRL: an open-source roguelike built specifically for that purpose. Inspired by DeepMind and OpenAI’s work, DelveRL offers a human-playable game with a structured API, deterministic simulation, and procedural generation—addressing a common integration hurdle. The included baseline agent achieves a median floor of 18, showcasing its potential.

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026
KDnuggets

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026

Evaluating AI coding agents demands rigorous benchmarks. In 2026, several open-source options will be essential for developers. Explore the top 10, including SWE-bench, Terminal-Bench, SlopCodeBench, and ProgramBench, alongside emerging contenders. These benchmarks offer critical insight into agent capabilities across diverse coding tasks. For deeper context on related AI research and development, see our discussion thread for EMNLP 2026 Notifications/Results. Discover how these tools empower informed decisions in the rapidly evolving landscape of AI-powered software engineering.

Machine Learning

How to make any Sparse Attention / KV Compression look good? [D] [R]

Navigating the complexities of Sparse Attention and KV Compression often involves presenting results that appear more impactful than they truly are. As detailed in a recent analysis by P. Nawrot, understanding these nuances—from carefully selected benchmarks to strategic prompt engineering—is crucial for accurate evaluation. This post explores common practices, like isolating contributions and leveraging aggregated metrics, that can inadvertently skew performance assessments.

Input 4-5x Reduction with sentence and keyword based trie on chat. [P]
Machine Learning

Input 4-5x Reduction with sentence and keyword based trie on chat. [P]

Users are reporting significant gains – up to a 4-5x reduction – leveraging a sentence and keyword-based trie for chat input retrieval. Currently, automatic budget selection faces challenges, occasionally retrieving excessive data despite promising accuracy near benchmark levels. We’re exploring algorithms beyond CELF to refine retrieval precision and enhance performance. This builds upon ongoing research into efficient attention mechanisms, as demonstrated in articles like "SSOG-Attention," which investigates scalable alternatives to SDPA. Discover how these innovations empower more effective data management.

Machine Learning

Did blatant AI Slop just win a 25K USD Deepmind / Kaggle Grand Prize? [D]

A recent DeepMind/Kaggle competition, "Measuring Progress Toward AGI," has sparked considerable debate following the announcement of its results. The 25,000 USD grand prize was awarded to a submission critiqued as presenting “nonsensical number generation” and questionable methodology. The work, intended to assess LLM reasoning through viewpoint comparison, appears to have been overlooked for critical review. Explore a deeper investigation of this outcome, detailing the methodology and data—a journey that may challenge conventional understanding.

Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]
Machine Learning

Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]

Researchers have introduced DABSN (Dynamic Adaptive Bias State Network), a novel recurrent language model architecture demonstrating promising results in reasoning, memory, and long-sequence tasks. The initial preprint and accompanying code—available in PyTorch, C++, and Triton—detail the architecture’s behavior and performance across benchmarks like MQAR and A5/60. Early language modeling experiments with a 24M parameter model have yielded unexpectedly strong results, prompting a second paper focused on scaling and long-context behavior. Collaboration is sought for independent reproduction, evaluation design, and access to larger GPU resources.