Evals

Evals on Beyond Market Intelligence: a running collection of 3 stories we have gathered and hand-picked because they are worth your time. Every post here touches on evals in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around evals, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
VentureBeat

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

Recent VentureBeat research reveals a concerning trend: 85% of companies that experienced an AI mistake are accelerating their move toward automated deployments, even as trust in automated evaluation rises. While automated checks are gaining traction, nearly half of surveyed enterprises still see test-approved AI features disappoint customers. This shift highlights a growing gap between evaluation confidence and real-world outcomes, prompting many to prioritize anomaly detection and issue resolution, as evidenced by the surging demand for platforms like Raindrop.ai.

Presentation: The Right 300 Tokens Beat 100k Noisy Ones: The Architecture of Context Engineering
InfoQ

Presentation: The Right 300 Tokens Beat 100k Noisy Ones: The Architecture of Context Engineering

Coding agents often falter, not due to insufficient context, but due to excessive and noisy input. In "The Right 300 Tokens Beat 100k Noisy Ones: The Architecture of Context Engineering," Baruch Sadogursky and Patrick Debois reveal why bloated context windows hinder performance and present practical fixes. Learn about lazy-loaded skills, versioned artifacts, and externalized memory—techniques to transform raw markdown into reliable agentic workflows.

Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026
VentureBeat

Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026

Expedia’s chief AI and data officer, Xavi Amatriain, is redefining product development, asserting that "evals are the new PRD." This shift prioritizes embedding security and design principles directly within evaluation processes, even before coding begins, leveraging AI-assisted code generation. Amatriain advocates for risk-calibrated governance layers, minimizing restrictive guardrails to maintain feedback loops and user agency, particularly emphasizing that users should retain the final click for transactions.