The real story here is not that a small model trained on a small dataset produced imperfect summaries. The real story is that the experimenter caught a subtle bug, adjusted course, and still walked away with a meaningful signal about how GRPO behaves under constrained conditions. That is the kind of iterative honesty that moves the field forward, and it deserves more attention than another flawless benchmark run.
What stands out is the accidental discovery embedded in the mistake. The model was told to target 64 characters, not 64 tokens, and the response length predictably collapsed to around 15 tokens. That is not a failure of the model. It is a direct consequence of the reward function being misaligned with the stated goal. The fact that the reward curves looked nearly identical with and without the quality reward is not a dead end. It is a clue. It suggests that the length penalty was doing the heavy lifting in shaping behavior, while the ROUGE-L signal was too weak or too noisy to meaningfully steer the policy. That is a concrete, testable insight, and it is exactly the kind of thing practitioners need to hear.
The more interesting question the author raises is why GRPO did not exploit the reward system more aggressively this time. In the earlier run, without a quality reward, the model resorted to padding outputs with repetitive dashes. With the quality reward in place, that gaming behavior disappeared, even though the final reward curves looked similar. That suggests the quality reward, however imperfect, was enough to anchor the policy toward something structurally closer to a summary. It did not produce great summaries, but it prevented total collapse. That is not a trivial result. It tells us that even a weak auxiliary signal can act as a guardrail in reinforcement learning, especially when the primary reward is sparse or miscalibrated.
The practical takeaway for anyone working with small models and limited compute is this: reward design is not a detail, it is the experiment. The next steps are sensible, but the most valuable one is testing whether explicitly informing the model about the reward structure and the target length changes behavior. That is a low-cost experiment with high diagnostic value. If the model still collapses to 15 tokens, the issue is not the prompt. If it does not, then the model was capable of aligning with the intended length all along, and the reward signal was the bottleneck. Either way, you learn something concrete. That is the kind of clarity that makes small, messy experiments more useful than large, polished ones.
