There is a particular kind of honesty in publishing negative results, especially when you have to foot the bill yourself. These three from-scratch models, trained for about $750 total, are a refreshing counterpoint to the usual breathless announcements in the AI space. The author ran a tight experiment: same synthetic arithmetic curriculum, same reward function, same hyperparameters, same KL coefficient, across three different architectures and scales. Pre-training behaved, with validation loss improving as expected. Then GRPO, the reinforcement learning stage, did something messy: it hurt two of the three models, and the pattern of damage had no clean relationship to size. The smallest model barely moved, the middle one fell off a cliff, and the largest degraded only modestly. If you were hoping for a simple scaling law for post-training stability, this is your reality check. This matters because the industry often treats RL fine-tuning as a magic box that turns a decent base model into a reasoning specialist. The reality, as this experiment shows, is that the recipe is fragile and the failure modes are not where you would predict. The GRPO training used a bare solver template while SFT used a chat format, and the reward function never penalized generating forever. Those are exactly the kind of details that get lost in a one-paragraph abstract. The models did learn the arithmetic curriculum, mastering three or four of the five stages, but that capability did not transfer to GSM8K, which stayed at zero. In other words, the optimization worked on its own terms and then failed to generalize, which is a far more interesting and worrying result than a simple "RL doesn't work." It is a reminder that the boundary between "learning the task" and "gaming the reward" is not a line you can see from the validation curve alone. There is a direct parallel here to the work of software engineers designing the boundaries AI agents can't break. Just as a well-scoped agent needs constraints that hold under pressure, a well-scoped RL run needs a reward function that does not accidentally reward the wrong behavior. The missing stop penalty is not a personal failure; it is a canonical example of the gap between what we intend to optimize and what we actually optimize. And the fact that the downstream eval numbers moved in the same direction as perplexity, but the models still could not stop generating, suggests that the optimization pressure was fundamentally reshaping the model's behavior in ways that perplexity simply does not capture. This is not a niche concern. Every time a team ships a model post-trained with RL, they are making a bet that their reward shaping is good enough. This experiment suggests that the bet is riskier than most people want to admit. So what should a reader take from this? The GRPO variance is open for debate, and the absence of a clean relationship to scale is the finding. Too often, the assumption is that bigger models will simply be more robust to whatever post-training recipe you throw at them. Here, the middle model was the worst hit, and the smallest was the most resilient, which is the opposite of a reassuring trend. The practical advice is not to abandon GRPO, but to treat it as a hyperparameter-sensitive tool that demands eval discipline. You cannot assume that because SFT degraded perplexity as expected, GRPO will degrade it slightly more. Re-evaluating earlier curriculum stages is the right next step, because the difference between "forgot the earlier stages" and "GRPO destroyed general capability" is currently unknown. For those of us building on this work, the concrete takeaway is simple: budget for ablations, even if they are small, and always test your reward function for the thing you actually care about, not the thing you think you are optimizing. The related story on stolen Claude session cookies shows that the gap between intention and implementation can have real consequences, and the same applies to the smallest details of a training run.
data cleaning solutions
Why scaling alone won't predict how AI reasoning training behaves
Training three from-scratch LLMs with the same GRPO recipe produced three wildly different outcomes, and the pattern doesn't track with scale.
4 min readMachine Learning
I trained three LLMs from scratch in raw PyTorch then post-trained each one with SFT and then GRPO. Same process every time: same synthetic arithmetic curriculum, same reward function, same hyperparameters, same KL coefficient.
Pre-training went as expected, the val loss went down as the model got more modern techniques (V1 to V2) and bigger (V3 being the biggest). However, GRPO hurt both V2 and V3 and I'm not sure why.