Expanding Chatterbox TTS to 8 Indian languages with just 1.4% of its parameters.

We're excited to announce the addition of eight Indian languages to Chatterbox-Multilingual, enhancing its voice capabilities for over 500 million speakers.

3 min readMachine Learning

The most impressive thing about this work is not the model itself, it is the discipline of the approach. Training just 7.8 million parameters out of 544 million to add eight languages, including four with entirely new scripts, is a deliberate act of restraint. The author did not chase a bigger model or a flashier dataset. They extended the tokenizer, warmed up new embeddings from phonetically related Devanagari characters, and let LoRA adapters do the heavy lifting. That is how you treat compute and data as the scarce resources they are. The result is a practical template for expanding speech technology to underserved languages without waiting for a lab with unlimited budget.

For anyone building in this space, the implications are immediate and concrete. You do not need to retrain a full model to add a language. You need a tokenizer extension, a smart initialization, and a training loop that respects what the base model already knows. The incremental training schedule is particularly worth studying. Adding languages one at a time, with weighted sampling, did not just prevent catastrophic forgetting. Hindi character error rate actually improved after seven other languages were added. That is not luck. That is a signal that the method is not merely preserving old knowledge, it is reinforcing it. The practical takeaway is that a small, well-structured fine-tune can outperform a sloppy full retrain, and it does so at a fraction of the cost.

That said, the results also draw a clear line between what is solved and what is not. Malayalam at a CER of 0.86 is not a minor blemish. It is a reminder that script complexity and data quality still dominate outcomes. The work is honest about this, and that honesty is a feature. The rest of the languages land in a range that is intelligible and, by the account of the work, natural sounding. But CER is not MOS. We do not yet know if these voices sound pleasant, expressive, or consistent across speakers. With only two speakers per language, the generalization question is wide open. These are not reasons to dismiss the work. They are reasons to treat it as a starting point, not a finish line.

The deeper point is that this is a replicable methodology, not a one-off hack. The author has shown that Brahmic scripts share enough phonetic structure that warm-starting from Devanagari works. That insight has shelf life. It can be applied to other related scripts and languages. The code is open, the model is on Hugging Face, and the training details are documented. That is what makes this worth paying attention to. Not the promise of a finished product, but the proof that a small team with a clear method can move the needle on a problem that has left half a billion speakers without representation. The next step is not more parameters. It is more languages, more data, and more people trying this and reporting back.

From Machine Learning

TL;DR: Fine-tuned Chatterbox-Multilingual (Resemble AI's open-source TTS) to support Telugu, Kannada, Bengali, Tamil, Malayalam, Marathi, Gujarati, and Hindi using LoRA adapters + tokenizer extension. Only 7.8M / 544M parameters trained. Model + audio samples available.

Chatterbox-Multilingual supports 23 languages with zero-shot voice cloning, but no Dravidian languages (Telugu, Kannada, Tamil, Malayalam) and limited Indo-Aryan coverage beyond Hindi. That's 500M+ speakers with no representation.

Read the original at Machine Learning