Reusable Evaluation Frameworks

From siloed AI evals to a unified framework for production-grade agentic products

Moving from ad-hoc AI agent evaluations to a unified, production-grade framework is the kind of transition most teams only talk about.

3 min readInfoQ
From siloed AI evals to a unified framework for production-grade agentic products

Elastic's move from siloed AI evaluations to a unified framework is the kind of maturity the agentic space desperately needs, and it deserves more attention than another flashy model release. Susan Chang's account of standardizing evals across Python data science workflows and TypeScript production code speaks directly to a problem every team building RAG or cybersecurity agents will eventually hit: your evaluations are only as trustworthy as the infrastructure holding them together. For readers still stitching together ad hoc tests, this is a signal that the discipline of evaluation is becoming a product problem, not just a research exercise.

The practical lesson here is about bridging worlds that rarely talk to each other. Chang's team balanced LLM-as-a-judge with deterministic rules, which is exactly the right instinct. Pure judgment calls from a model are seductive but brittle; pure rules miss the nuance of domain-specific workloads. By pairing them, Elastic gets the flexibility of AI-based scoring with the guardrails of reproducible logic. That hybrid approach should resonate with anyone who has watched a great demo fail in production because the evaluation set never captured the messy reality of user queries. It also connects to the broader challenge of testing retrieval systems, a topic we explored in Uncover Retrieval Weaknesses: Test Your RAG Pipeline Now, where small adversarial sets can expose failures your standard metrics miss.

What stands out most is the emphasis on deep tracing to catch regressions across complex workloads while preserving domain context. This is not about building a bigger test suite; it's about making evaluations observable enough to trust. If you cannot trace why an agent failed, you cannot fix it, and you certainly cannot scale it. Elastic's approach suggests that the next frontier is not more eval data but better instrumentation, a point that pairs naturally with Bridging Retrieval and Action: A New Approach to AI Tasks, where connecting retrieval and action explicitly demands the kind of visibility Chang describes. Teams that skip this tracing layer are flying blind, no matter how clever their prompts are.

The takeaway is concrete: unify your evaluation strategy before your agent complexity outpaces your ability to debug it. Start by mapping which of your current evals are deterministic, which rely on model judgment, and where you lack tracing entirely. If you are building production agents, this is not an infrastructure nicety; it is the difference between a system you can improve and one you can only hope works. The open question worth watching is how Elastic's framework handles evaluation drift over time, since domain context is not static. But for now, the direction is clear, and it is one more reason to treat evaluation as a first-class engineering discipline rather than an afterthought.

From InfoQ

Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context.

Read the original at InfoQ