The final answer is the least interesting thing about an AI agent. That uncomfortable truth is exactly what one developer discovered while running local agents with Ollama and LangChain, and it deserves far more attention than it's getting. You can get a perfectly correct final output while the agent is doing absolute nonsense internally, calling the wrong tool first and then recovering, using tools it never needed, looping several times before converging, or even creeping dangerously close to actions it should never take. If you're only checking the output, all of that passes without a second thought.
For anyone building or deploying local agents, this has immediate practical consequences. Two agents can both summarize a document correctly. One does it in two clean steps: read, then summarize. The other does read, search, read again, summarize, retry. Same result. But one is clearly more efficient, less risky, and more trustworthy. If you're not looking at the trace, you treat them as equal. That's not just a quality gap, it's a blind spot that can hide serious problems. Loops waste compute. Unnecessary tool calls increase latency and surface area for errors. And the wrong tool used even once, even if corrected, can mean a security boundary was crossed or data was mishandled.
The developer who ran into this built a small local eval setup that checks tool usage against expected and forbidden lists, penalizes loops and extra steps, and runs entirely locally with Ollama as the judge. That's a practical starting point, but it also highlights how little conversation exists around trace-level evaluation on the local side. Most eval setups still focus heavily on final answers, or they assume you're fine sending data to an external API for judgment. For local agents, that approach misses the point. The process is where all the signal lives. Metrics like tool efficiency, loop detection, and reasoning coherence matter more than whether the last token is correct.
So here is the concrete takeaway: if you are evaluating agents by final output alone, you are not evaluating agents, you are evaluating answers. The agent is the process, and the process needs its own metrics. Start tracking tool selection, step count, loop frequency, and reasoning quality. Build local validation that matches your local deployment. The developer's repository is a useful starting point, but the real work is making trace evaluation a standard practice, not a hack.