rows.com

Small model, sharp summaries: refining GRPO with a simple fix.

In a recent experiment, I trained a Qwen2.5-0.5B-Instruct model on the smoltldr Reddit post summarization dataset, targeting concise outputs of 64 tokens. However, I mistakenly set the limit to 64 characters, resulting…

3 min readMachine Learning
Small model, sharp summaries: refining GRPO with a simple fix.
Trained a Qwen2.5-0.5B-Instruct bf16 model on Reddit post summarization task with GRPO [P]

The real story here is not that a small model trained on a small dataset produced imperfect summaries. The real story is that the experimenter caught a subtle bug, adjusted course, and still walked away with a meaningful signal about how GRPO behaves under constrained conditions. That is the kind of iterative honesty that moves the field forward, and it deserves more attention than another flawless benchmark run.

What stands out is the accidental discovery embedded in the mistake. The model was told to target 64 characters, not 64 tokens, and the response length predictably collapsed to around 15 tokens. That is not a failure of the model. It is a direct consequence of the reward function being misaligned with the stated goal. The fact that the reward curves looked nearly identical with and without the quality reward is not a dead end. It is a clue. It suggests that the length penalty was doing the heavy lifting in shaping behavior, while the ROUGE-L signal was too weak or too noisy to meaningfully steer the policy. That is a concrete, testable insight, and it is exactly the kind of thing practitioners need to hear.

The more interesting question the author raises is why GRPO did not exploit the reward system more aggressively this time. In the earlier run, without a quality reward, the model resorted to padding outputs with repetitive dashes. With the quality reward in place, that gaming behavior disappeared, even though the final reward curves looked similar. That suggests the quality reward, however imperfect, was enough to anchor the policy toward something structurally closer to a summary. It did not produce great summaries, but it prevented total collapse. That is not a trivial result. It tells us that even a weak auxiliary signal can act as a guardrail in reinforcement learning, especially when the primary reward is sparse or miscalibrated.

The practical takeaway for anyone working with small models and limited compute is this: reward design is not a detail, it is the experiment. The next steps are sensible, but the most valuable one is testing whether explicitly informing the model about the reward structure and the target length changes behavior. That is a low-cost experiment with high diagnostic value. If the model still collapses to 15 tokens, the issue is not the prompt. If it does not, then the model was capable of aligning with the intended length all along, and the reward signal was the bottleneck. Either way, you learn something concrete. That is the kind of clarity that makes small, messy experiments more useful than large, polished ones.

From Machine Learning

So, a few days back I shared a post where I trained a tiny Qwen2.5-0.5B-Instruct model on smoltldr (reddit post summarization dataset of 2k rows), to output summaries of about 64 max length using RLVR with GRPO .

Hence the charts showed a sharp decline and convergence towards a response length of on and off 15 tokens.

Read the original at Machine Learning