Tokenization is the invisible engine of every language model you have ever used, and it has been treated like a plumbing problem for too long. A new survey from 32 researchers finally gives this hidden layer the attention it deserves, and the findings should change how we think about model performance, multilingual support, and even security. This is not an academic curiosity, it is a practical wake-up call for anyone building or evaluating AI systems today.
The survey covers everything from algorithms to adversarial attacks, but what matters most to practitioners is the concrete evidence that tokenization choices ripple outward in ways we have underestimated. For example, the way a model breaks text into tokens directly affects how well it handles languages beyond English, how reliably it follows constrained generation instructions, and even how vulnerable it is to manipulation. This connects directly to questions we have explored before, such as Rethinking the compute demands behind LLM post-training research, where the hidden costs of standard practices often go unexamined. Similarly, the survey's discussion of token healing and constrained generation echoes themes from Beyond the Decoder: Choosing When to Generate, Not Just Decide, the idea that the architecture of generation matters as much as the output itself. And when the survey explores replacing tokenizers with latent or visual tokenization, it aligns with the insight from Your camera roll already knows more about your life than your inbox does: that different data modalities demand different processing strategies.
Our take is straightforward: the industry has been shipping models without fully understanding one of their core components. Tokenization is not a solved problem to be inherited from the nearest open-source repository, it is a design decision with measurable consequences. The survey documents how tokenizer design affects everything from vocabulary efficiency to cross-lingual parity, and it identifies gaps in evaluation that mean many teams are likely optimizing for the wrong metrics. For anyone deploying models in production, the specific takeaway is this: audit your tokenizer the same way you audit your training data. The survey shows that tokenization artifacts can produce systematic biases that no amount of fine-tuning will fix.
The most provocative section of the survey looks at what might replace tokenizers entirely, latent tokenization and visual tokenization approaches that bypass the text-splitting step. This is where the field could move in the next few years, and it raises an open question that every AI team should start thinking about now: if your model's vocabulary is an artifact of engineering convenience rather than linguistic necessity, how much are you leaving on the table? The answer, according to the 32 researchers who compiled this work, is probably more than you think.