LLM Agents

EvoUndo gives AI agents the power to evolve safely without losing control.

Self-modifying agents are risky.

4 min readMachine Learning

The EvoUndo paper lands at a moment when the industry is busy celebrating agents that can rewrite their own prompts or swap out tools mid-task. That capability is impressive, but it carries a quiet risk that most discussions gloss over: what happens when a successful mutation leaves a permanent mark that cannot be undone in a different context? The authors found 197 such failures out of 600 tasks, and conventional repair strategies recovered none of them. That is not a rounding error. That is a warning sign that we have been optimizing for capability while ignoring recoverability, and EvoUndo makes a compelling case that these two goals must be designed together.

What stands out is the distinction between grounding and expressivity. When the original recovery language was sufficient, adding exact state-address diagnostics improved recovery from zero to 38 out of 48 failures. That is a meaningful jump, but it also reveals that the bottleneck was not just about having the right words. It was about being able to point at the precise state that needed to change. On the other side, extending the recovery language allowed the system to handle 142 out of 143 failures in the richer stratum. The lesson is practical: if you want agents that can evolve safely, you need both precise state awareness and a vocabulary broad enough to describe the undo. Iterative prompting alone will not get you there, and the paper's replication on a different backbone confirms that this is not a quirk of one model.

We have covered adjacent territory in our recent coverage of Unlock 2D Rotations: Exploring the Power of Complex Kimi Delta Attention, where expressivity in attention mechanisms turned out to be the deciding factor in what a model could represent. The parallel is direct. EvoUndo shows that recovery language expressivity plays a similar role for undo operations, and that constraining it too tightly creates hard ceilings on what can be reversed. Likewise, our look at Presentation: Context Engineering at LinkedIn: How We Built an Organizational Context Layer for AI Agents with MCP highlights how much of agent reliability depends on the surrounding infrastructure rather than the model alone. EvoUndo reinforces that point, but with a sharper edge: the infrastructure must include explicit verification and recovery mechanisms, not just context plumbing.

The most interesting result is the negative interaction on the primary backbone, where adding exact-address diagnostics to the richer language actually reduced recovery from 99.3 percent to 93.0 percent. That is counterintuitive, and it suggests that more information is not always better when it is fed into a system that was not designed to handle it. The Qwen replication did not show this effect, so this is model-dependent behavior that we do not yet understand. That is the open question worth watching: under what conditions does additional diagnostic detail confuse rather than clarify? If we cannot answer that, we risk building agents that are powerful in controlled settings but brittle in the messy reality of production use.

For anyone building agent harnesses today, the takeaway is direct. Do not assume that a model can reliably undo its own modifications. Build verification into the loop from the start, and treat recovery language as a first-class design constraint rather than an afterthought. The paper does not offer a silver bullet, but it gives us a concrete framework for measuring recoverability, and that is a step forward we should all take seriously. Watch for the next round of work on how diagnostic grounding interacts with richer languages across different model families, because that is where the real answers will come from.

From Machine Learning

LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created.

Read the original at Machine Learning