Self-Healing Agents: When Your System Routes Around Its Own Bugs

I'm excited to share insights from my recent experience with Curragh, a local instance of my project, Springdrift.

3 min readMachine Learning

The most interesting thing in this story isn't that Curragh detected a missing agent. Any system with a decent error log can do that. What matters is that it diagnosed the root cause, articulated the impact, and then chose a workaround without being prompted to run a single diagnostic tool. That is a different kind of agency, and it deserves attention.

The key architectural choice is the passive sensorium. Curragh didn't have to ask itself, "What's wrong?" It already had the live agent roster, the pipeline definition, and the error message in its context window. That meant the problem was visible the same way a scratch on a desk is visible: no inspection required. In a session-bounded agent, you would have to explicitly request a diagnostic, or you would just watch it retry the same failing call until you stepped in. Curragh instead treated the environment as something to read continuously, not something to interrogate on demand.

For practitioners, this flips the cost model of self-correction. Tool-based introspection is expensive and brittle. It assumes you know what to ask for before you know what's broken. Passive context removes that assumption. The agent sees the state it needs to see, and when something is off, it can reason about it in the same cycle where it notices. That is closer to how a human engineer works, not because we are mystical, but because we don't run a separate "diagnose the system" function every time we glance at a dashboard. We just look.

The anthropomorphic framing is a trap, but a useful one. Curragh is not conscious, and it did not "decide" to be helpful. But the behavior is real, and it emerged from a concrete design choice: append-only memory, OTP supervision, and a structured self-state block injected each cycle. Those are not vague aspirations. They are specific mechanisms that produce observable outcomes. If you are building agents that need to recover from their own bugs, you should study this example. Then you should ask yourself why your agent still needs to be asked what's wrong.

From Machine Learning

Sharing my project Springdrift for feedback. It's a persistent runtime for LLM agents, append-only memory, OTP supervision, a structured self-state block injected each cycle instead of requiring tool calls for introspection(sensorium). I've been running a local instance (Curragh) for a few weeks and wanted to share a recent episode. It's not the strongest example (check the paper) but it's recent and undocumented. Fair warning: I'll describe what Curragh did in somewhat anthropomorphic terms because that's how I think about it when I'm working with it. The paper is stricter.

Read the original at Machine Learning