Reinforcement Learning

Teaching PPO to Track the Ball Instead of Memorizing Scores

After 124 failed PPO experiments on Atari Breakout, Mikey Harrell found something worth celebrating: every model defaulted to a memorized script.

4 min readMachine Learning

Six months of experimentation, 124 failed attempts, and the breakthrough came down to three lines of code. That is the story Mike Harrell tells about his journey to make PPO play Atari Breakout like a human instead of a memorizing machine. For anyone who has spent time training reinforcement learning agents, the frustration he describes will feel immediate and familiar. The models kept finding scripts, not strategies. They learned to game the system, not to play the game. And no amount of environment tweaking, sticky actions, or entropy tuning could push them off that path. The default optimum in reinforcement learning is a shortcut, not understanding.

What makes this worth pausing over is not the technical fix itself, though the proximity reward is elegant. It is what the failure mode reveals about how we approach complex systems. Harrell tried everything he could think of to make the environment harder to memorize, and PPO just adapted. Timing-robust scripts. Layout-conditioned scripts. Noise-tolerant scripts. Every time, the policy found a way to game the reward structure without ever tracking the ball. That is not a bug in PPO. That is the objective function doing exactly what it was told. The problem was that the objective was points, not play. This is a lesson that extends far beyond Breakout. Whether you are tuning a model or building a tool, the reward you define is the behavior you will get. If you want reactive intelligence, you have to reward the reaction itself, not just the outcome.

The shift Harrell landed on is a small but meaningful reframe. Instead of trying to block memorization, he changed what the optimum looks like. A tiny per-frame bonus for paddle proximity to the ball during descent made tracking the highest-reward strategy. The script still gets occasional incidental bonuses, but it cannot compete with a policy that actively follows the ball. The behavior transfers to clean evaluation, which is the part that matters. This is not about adding complexity. It is about aligning the incentive structure with the capability you actually want. That is a principle that scales. In the same way Kubernetes 1.37 Released: Stable Metrics API and Rootless Kubelet in Beta signals that infrastructure work is moving toward more stable, user-facing abstractions, Harrell's approach points to a simpler kind of progress: defining success in terms of behavior, not just output.

What we would tell a reader asking about this is to look past the Breakout demo. The Split-Watcher tool is compelling, and watching the agent track the ball across custom brick configurations is a small thrill. But the real insight is that 123 failures are not wasted effort. They are a map of the optimization landscape. Harrell did not just find a solution. He documented the shape of the problem, and that documentation is arguably as valuable as the fix. For anyone training models on real-world tasks, the takeaway is direct: if your agent is memorizing, do not just make the environment harder. Ask what behavior you are actually rewarding. You might find, as Harrell did, that a few lines of reward shaping do more than a hundred environment tweaks. The open question is whether this proximity-reward trick generalizes beyond games. We suspect it does, but that is the next experiment. And we will be watching to see what else three lines of code can unlock.

From Machine Learning

Six months ago I started experimenting with PPO and Breakout as a way to learn about Machine Learning and Reinforcement Learning. After a few experiuments just trying to get high scores, it bothered me that everything was a "memorized" script rather than reactive play, like a human would play. Thus began my journey to try and convince PPO to actually track the ball instead of focusing on scoring points. I read a lot of articles and tried a lot of things. After 124 PPO experiments on Atari Breakout, I found that every single model, across sticky actions, cursor wrappers, entropy…

Read the original at Machine Learning