Sona

One transformer model replaced 15 specialized generators in Yandex Music's A/B test

One transformer replaced 15 specialized generators in Yandex Music's production stack.

3 min readMachine Learning

The core idea behind Sona, one transformer model replacing fifteen specialized generators in Yandex Music's production stack, is exactly the kind of practical simplification that makes us stop and take notice. This isn't a theoretical paper about what might be possible in five years. It's an A/B test result showing a +4.53% lift in Active Users and +6.30% increase in Total Listening Time on smart speakers, with statistical significance at p < 0.01. That is a concrete outcome, and it tells us something important about where recommendation systems are heading.

What makes this worth exploring is the engineering trade-off the team made to get there. Full attention over 8,192 events is expensive, so they introduced History Compression: splitting the user's history into an older block of 6,144 events and a recent block of 2,048, then using cross-attention and a single full-history self-attention layer to let the two blocks talk to each other. The result is roughly half the inference cost while retaining most of the quality. This is a smart, pragmatic approach to a problem that often gets solved by throwing more hardware at it. It echoes the thinking behind Build leaner AI apps with smarter token compression techniques, where the focus is on reducing cost without sacrificing output quality. The Yandex team didn't just prove that a single transformer can work, they proved it can work efficiently enough to run in production.

The practical takeaway for anyone building recommendation systems is direct: you may not need a sprawling pipeline of candidate generators, pre-ranking models, and ranking models with hundreds of features. A single, well-designed transformer can handle the entire task, from reading user history to generating candidates via beam search to scoring them, all from one encoder pass per request. That is not a marginal improvement in architecture; it is a fundamentally different approach to how you structure a recommender. And it works. The fact that catalog coverage was lower than the production stack is an honest signal worth watching, the team is investigating why, and that detail matters. It suggests the model may be optimizing for engagement over diversity, a tension that every recommendation system eventually faces. The long-term A/B test now underway will tell us whether that trade-off holds or can be addressed.

We are also reminded, by the paper's own reporting, that this kind of simplification did not come from nowhere. The LLM community has been showing for years that end-to-end generative models can absorb work previously split across specialized components. What Sona does is carry that recipe into music recommendation and prove it works in a live A/B test. This is not a claim about replacing all specialized systems overnight. It is a concrete demonstration that How language models learn to copy context with hash tables has real, practical applications beyond text. The question that remains open is whether this approach generalizes to other domains, video, e-commerce, news, where user histories are longer and the cost of full attention becomes prohibitive. Yandex's History Compression technique offers one path forward. We will be watching the long-term results closely.

From Machine Learning

Our production recommender at Yandex Music has 15+ candidate generators feeding pre-ranking and ranking models with hundreds of features. LLMs showed that one end-to-end model can take over work that used to be split across specialized components, and single-model generative recommenders have carried that recipe into production. We set out to explore what a single-model recommender could do in music. The result is Sona, one transformer that replaced all of it in an A/B test. It hasn't shipped to full traffic yet.

Read the original at Machine Learning