The real gap in AI isn't pretraining: it's who gets to do the reinforcement learning.

The dominance of large machine learning labs in the development of widely-used models, such as GPT and Claude, can be attributed to several factors beyond just the costly pretraining.

3 min readMachine Learning

The conversation about why models like GPT and Claude dominate real-world usage often defaults to the cost of pretraining. But that explanation misses the real story. The expensive pretraining is already done, open-source models like Kimi and DeepSeek have proven that. What separates the leaders from the rest is the reinforcement learning that happens after, and that's where the true bottleneck lies. It's not about who can afford to build the base model; it's about who can refine it into something people actually want to use.

This matters because it reframes the entire competitive landscape. If RL is the differentiator, then the barrier to entry isn't just financial, it's expertise, data quality, and the ability to iterate quickly. Smaller labs and independent teams have access to the same pretrained foundations, but they often lack the infrastructure to run effective RL pipelines. That's not a technology gap; it's an operational one. The labs that excel aren't necessarily the ones with the most compute. They're the ones that have figured out how to turn a good base model into a great product through careful, continuous feedback loops.

For the reader, this is both a warning and an opportunity. If you're building on top of open-source models, the raw capability is already there. What you need to invest in is the layer above it, the data curation, the reward modeling, the human feedback systems. That's where the value is being created now. It's not enough to fine-tune; you need to think about how your model learns from interaction, how it aligns with user intent, and how it improves over time. The labs that dominate aren't doing something you can't do; they're just doing it more deliberately and with better feedback mechanisms.

The practical takeaway is simple: stop worrying about who has the biggest pretraining budget. Start focusing on how you're going to do the RL. That's where the next wave of differentiation will come from, and it's far more accessible than you might think. The tools and models are already out there. The question is whether you're prepared to put in the work to make them sing.

From Machine Learning

I’m trying to understand why models from major labs (GPT, Claude, etc.) dominate real-world usage? You might say it's due to the expensive pretraining compute budge, but there already exists many pretrained open-source models at the same scale (e.g., Kimi).

Read the original at Machine Learning