agent evaluation

3 stories filed under agent evaluation on Beyond Market Intelligence. The newest of them: “QCon London 2027 Maps 15 Tracks for Smarter AI and Engineering at Scale”, “From Spreadsheets to AI Assistants: A Principal Scientist's Journey”, and “Evaluate AI Coding Agents with These Top Open-Source Benchmarks”. QCon London 2027 is shaping up to be a serious map of what's next. James Gung, a principal applied scientist at AWS, is opening the floor to questions about building AI services like Amazon Bedrock and Lex. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work… The list below is every agent evaluation story on Beyond Market Intelligence, newest first.

QCon London 2027 Maps 15 Tracks for Smarter AI and Engineering at Scale
InfoQ

QCon London 2027 Maps 15 Tracks for Smarter AI and Engineering at Scale

QCon London 2027 is shaping up to be a serious map of what's next. With 15 tracks spanning agent evaluation, guardrails, AI-era architecture, and the messy reality of distributed-system debugging, the program doesn't just celebrate scale; it confronts it. For teams wrestling with modern data platforms or Staff+ leadership, this is a practical guide to building smarter. And if you're curious how routing fits into that picture, our related piece on the Jev approach offers a useful counterpoint.

From Spreadsheets to AI Assistants: A Principal Scientist's Journey
Machine Learning

From Spreadsheets to AI Assistants: A Principal Scientist's Journey

James Gung, a principal applied scientist at AWS, is opening the floor to questions about building AI services like Amazon Bedrock and Lex. His work spans agent evaluation and conversation simulation, and he's offering a candid look at the day-to-day reality behind those systems. It's a useful chance to separate the substance from the job-title noise. For anyone navigating similar roles, his perspective on interviews and career paths is worth exploring, especially when the AI job market feels muddled.

Evaluate AI Coding Agents with These Top Open-Source Benchmarks
KDnuggets

Evaluate AI Coding Agents with These Top Open-Source Benchmarks

Choosing the right benchmark is the first step toward building an AI coding agent you can actually trust. In 2026, the landscape has matured well beyond simple unit tests, with tools like SWE-bench, Terminal-Bench, and SlopCodeBench offering distinct lenses on real-world performance. We break down the top 10 open-source options to help you navigate the noise. If you are questioning how these models handle nuance, our piece on talking to an AI clone offers a timely perspective on the human side of the equation.