LLMs

Demystifying the math behind the algorithms powering modern LLMs

If you've been following the technical reports from Kimi, DeepSeek, Qwen, and GLM, you've likely noticed how much on-policy distillation and GRPO-style algorithms now power the frontier.

4 min readMachine Learning

The clearest signal in the recent wave of LLM tech reports isn't a new benchmark or a clever prompt. It's the quiet, consistent emphasis on reinforcement learning, specifically GRPO-style algorithms and on-policy distillation. When a practitioner like John Olafenwa takes the time to break down the maths and code behind these methods, he's pointing at something important: the frontier isn't just about bigger datasets anymore. It's about how models learn from feedback loops, and that shift is happening in plain sight. His deep dive is a useful reminder that the gap between pretraining and supervised fine-tuning is now being bridged by techniques that most of us haven't fully mapped out yet.

For anyone who has been following the field, this feels like a natural evolution. We've spent years talking about next-token prediction and loss curves, but the real leverage now sits in how a model refines its own policy after initial training. The fact that Kimi, DeepSeek, Qwen, and GLM are all leaning on these approaches tells you that this isn't a niche research interest. It's becoming the default operating system for frontier models. And yet, most practical guides still treat RL as an afterthought, a footnote in the training pipeline. That's a mismatch. If you're trying to understand why a model reasons better or follows instructions more reliably, you need to look at the RL layer, not just the base architecture. This is exactly why we've been tracking related ground in our piece on Unlock LLM Training: A Practical Guide to Distributed Algorithms. The fundamentals of distributed systems and RL are converging, and ignoring that connection leaves you with an incomplete mental model.

What's striking about Olafenwa's offer to answer questions is how accessible he's trying to make this. He's not gatekeeping the maths behind a paywall or assuming you already have a PhD in optimization. That matters, because the barrier to entry in this field isn't just compute. It's conceptual. If you're a developer or a data scientist who came up through traditional software engineering, the jump from supervised fine-tuning to policy-based training can feel abrupt. The job market is already shifting too, as we've seen in the changing expectations for AI and ML roles, where employers are asking for a blend of software engineering skills and a working understanding of training dynamics. The days of treating model training as a black box are ending. You don't need to be able to implement GRPO from scratch to be effective, but you do need to know what it is, why it works, and when to reach for it.

Our take is straightforward: this is the direction the field is heading, and the sooner you get comfortable with the mechanics, the better positioned you'll be. The open question that remains is how far on-policy methods can take us. We're seeing early evidence that they help with reasoning and alignment, but we haven't yet hit the ceiling of what's possible. Watch for the next wave of research to push on this. And if you're curious about the practical side, Olafenwa's video is a solid starting point, but don't stop there. Pair it with a look at how Neurosurgery Match Requirements Highlight Growing Pressure on Medical Students illustrates the broader trend of specialized knowledge becoming a competitive differentiator. The same logic applies here: the more you understand the training process itself, the less you're at the mercy of someone else's API. That's the edge worth building.

From Machine Learning

Hi everyone, if you have been reading the tech reports of Kimi, DS, Qwen and GLM, you will realize how much on policy distillation and GRPO style algorithms power the frontier. I thought it will be quite beneficial to do a deep dive explaining the maths and code behind this algorithms and how they connect to pretraining and supervised fine tuning.

I have published a deep dive on this topics here

Read the original at Machine Learning