GRPO

Recreating AI alignment results on consumer hardware reveals subtle training challenges.

A GRPO run that moves a trait only +2.4 points when you need +15 isn't a slow start; it's a signal that something structural is off. Feedback aligns with what many small-scale RL practitioners suspect: prompt diversity…

4 min readMachine Learning

There's a particular kind of honesty that comes from sharing a failed reproduction in public, and that's exactly what makes this post worth sitting with. The researcher behind this GRPO run isn't asking for applause. They're asking for help, and they've done the hard, unglamorous work of ruling out reward hacking, memorization, dead gradients, and even a real confound in their completion-length cap. That's not a failure of rigor. That's a masterclass in it. The result is a trait score that moved only +2.4 points on a 100-point scale, when the target was closer to +15. On the surface, that looks like a stall. But look closer, and the real story is about how much we still don't know about installing stylistic traits through reinforcement learning, even with a mechanically healthy training loop.

The diagnosis, confirmed by an author of the original paper, is that 20 distinct trait prompts is simply too few, and that per-example prescriptive rubrics likely matter more than a single global rubric. That aligns with what we've seen in Explore How Verifiable Rewards Empower Small Language Models, where the reward function does as much heavy lifting as the model itself. If the reward signal is too coarse, the model has no way to learn the nuance of what "consistent" or "traditional" actually means in a given context. But here's the uncomfortable question: why would prompt count and rubric specificity matter so much more for a stylistic trait than for a task-like trait? The answer might be that tasks have a single correct answer, while traits are relational. A model can't just learn to output the right thing. It has to learn when to suppress its default behavior, which is a far more delicate optimization problem.

This is where the post connects to a broader tension in the field. We're seeing a lot of enthusiasm for reinforcement learning as a general-purpose tool for shaping model behavior, and Is Reinforcement Learning Really Needed for Jev's Spreadsheet AI? asks a version of that question from the product side. But this reproduction attempt shows that RL isn't a magic dial. It's a surgical instrument, and when you're working with a 7B model on a single 3090, you're operating at roughly a hundred-thousandth of the compute the original paper used. The fact that the researcher got any movement at all, given that constraint, is a signal that the method has traction. The fact that it's not enough is a signal that we need to rethink what "small scale" means for trait installation, not just for persistence.

Here's the takeaway we'd offer anyone trying to walk this path: don't chase the persistence result yet. The install is the gate, and the gate swings on data diversity and reward granularity, not on raw compute. The next step is obvious, and it's the right one. They need to triple or quadruple the number of distinct trait prompts, and they need to write specific, per-prompt rubrics that tell the judge exactly what "consistent" looks like in that scenario. That's not glamorous. It's not even particularly novel. But it's the difference between a model that performs a trait and a model that just performs. Watch for their update. If the +2.4 jumps to +15 with those changes, we'll have learned something concrete about how to shape personality in small models. If it doesn't, we'll have learned something even more valuable about the limits of RL for stylistic control. Either way, this is the kind of open, methodical work that moves the field forward.

From Machine Learning

TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 points (95% CI [+0.2, +4.8]) when I need ~+15. Training is mechanically healthy and I’ve ruled out the obvious culprits. Looking for advice from people who’ve done small-scale RLHF/GRPO trait or persona installation.

Read the original at Machine Learning