Harness Training

Train a harness once, then empower any model on any task

Training a harness once, then swapping in any task LLM, is the kind of reframe that turns a clever hack into a real discipline.

4 min readMachine Learning

Most people treat an AI harness like a fixed tool, something you bolt onto a model and hope for the best. This project from Henry Pan reframes that assumption, and the reframing is worth sitting with. Instead of endlessly prompting or fine-tuning the task model, he proposes training the harness itself, once, against a frozen task LLM and a specific environment. After that, the harness is frozen and then evaluated against entirely different models and tasks. The code example is deceptively simple: a `Trainer` class, a `StrictPareto()` criterion, a `GreedyMonotonic()` optimizer, and a loop that records baseline-versus-candidate verdicts before deciding whether to commit or reject a change. That is the whole loop. It is the kind of clarity that makes you wonder why more projects in this space are not built this way.

The practical payoff here is not just academic. Pan reports that a harness trained on SWE-Bench tasks transfers to solve Terminal-Bench tasks, which is a concrete demonstration of capability transfer rather than mere memorization. That is a big deal for anyone who has watched agentic systems fail the moment you swap the environment. The project also leans on determinism as a core lesson, which is refreshingly honest given how much noise dominates the agentic AI conversation. He openly notes that his initial version was missing it, and that admission is more useful than most polished demos. For our readers building their own AI portfolio projects, this is a reminder that the hard part is not writing a loop, it is designing the reward signal and the rollback mechanism so the loop actually learns.

What stands out is the modularity. The framework is built like a PyTorch trainer, but instead of tensors, you are working with git commits and API calls. That abstraction is what makes it model-agnostic and task-agnostic, and it is also what makes it practical. You can swap out the task LLM without retraining the harness. You can point it at a new environment and see if the frozen policy holds. This is the kind of architectural thinking that moves the field forward, not because it introduces a flashy new model, but because it treats the harness as a first-class citizen worthy of optimization. It echoes the DSP-inspired semantic vocoder approach in that both projects find a clever abstraction to make a messy problem tractable, one in signal processing, one in agent orchestration.

The open question, and the one we would push Pan on, is evaluation depth. The results are promising, but beating Terminal-Bench with a frozen harness is not the same as beating a real-world production workload with noisy inputs and shifting objectives. The framework is extensible, and that is its strength, but extensibility also means the user is responsible for defining what "better" means beyond a Pareto frontier. For readers, the takeaway is direct: if you are building agents, stop tuning the model and start training the harness. The specific project is worth cloning and studying, not because it is finished, but because it shows a viable path forward. And the next time someone tells you agents are too unpredictable to trust, point them to a loop that commits or rejects based on evidence, not vibes.

From Machine Learning

I worked on this project (https://github.com/workofart/harness-training) for the past few months to reframe "Agent-driven Self-improving Harness" to "Harness Training".

The idea is simple, the harness is trained once with a frozen task LLM against a given task environment. Then you can then swap out the task LLM to any model and evaluate the "frozen trained harness" with any task LLM on any new task environment.

Read the original at Machine Learning