generative AI for data analysis

Discover how open-weight TTS brings expressive narration to your local machine.

TontaubeV1 takes a character-level approach to text-to-speech, and that decision is worth pausing on.

4 min readMachine Learning
Discover how open-weight TTS brings expressive narration to your local machine.
We released TontaubeV1, a character-level TTS model for long-form generation [P]

Open-weight speech models tend to chase scale and raw fluency, but the team behind TontaubeV1 has made a pair of design choices that feel genuinely instructive rather than merely incremental. Their decision to force character-level tokenization on a Qwen3 backbone, instead of defaulting to the model's native BPE tokenizer, runs counter to the prevailing trend in LLM-based TTS. Most projects treat the backbone tokenizer as fixed infrastructure, but the developers found that rare token combinations, especially around special characters, pushed the model out of distribution. That is a practical insight for anyone who has tried to build on top of a general-purpose language model and discovered that fine-tuning on audio does not automatically fix text-level fragility. It also connects to a broader conversation about how we repurpose general models for specialized tasks, much like the questions raised around [ProgramAsWeights: compile English function descriptions into neural programs that run locally [R]](/post/programasweights-compile-english-function-descriptions-into-cmu91wk4f04ghngmmxodderg), where the core issue is also about what a model loses or retains when you change its expected input format.

The chunking and positional scheme is the second quiet rebellion, and it is arguably the more important one for long-form narration. TontaubeV1 does not just split text into chunks and hope for the best; it assigns logical position IDs that keep text and audio on the same timeline, while using physical sequence order to control what the model can attend to. That distinction matters because it solves a real pain point in TTS: the model needs to know how far apart two sounds are in time, but it also needs a bounded context window for arbitrarily long passages. The 25-character reserved buffer at each boundary is a small detail, but it is exactly the kind of hack that separates a research demo from something that can sustain a 400-passage audiobook without drifting into incoherence. And the fact that they are already thinking about overlapping DualCodec windows for streaming shows a focus on deployment reality, not just benchmark scores. This is the kind of engineering that reminds us why open-weight efforts matter, just as the Databricks Acquires Row Zero, Signaling Future AI-Native Spreadsheet Growth story highlights how quickly AI-native interfaces are moving into productivity tools that were once considered static.

So what should a reader actually take from this? The benchmark results against ElevenLabs Flash v2.5 are promising, but the authors themselves caution that human listening tests are the gold standard, and they have not run one yet. That is the right caveat, and it should temper any impulse to declare a new leader. What is more valuable is the transparency around failure modes: the technical report openly discusses where BPE tokenization caused out-of-distribution behavior, and the choice to go character-level is presented as a trade-off, not a magic bullet. For practitioners, the concrete lesson is that your tokenizer is not a neutral preprocessor; it is a prior over what kinds of sequences your model can handle. If you are building a TTS pipeline, forcing character-level inputs on a model that was not trained that way might feel like a step backward, but it can actually make the mapping from text to sound more direct. The open question to watch is whether the 24 GB VRAM requirement comes down meaningfully with the promised quantized releases, because that will determine whether this becomes a tool for tinkerers or just for labs with deep pockets. That is the detail to follow.

From Machine Learning

My brother and I just released TontaubeV1, a 2.9B-parameter open-weight TTS model focused on expressive speech, long-form generation/narration, and low-latency local inference. It is primarily aimed at English and German and supports zero-shot voice cloning from up to one minute of reference audio. It builds on DualCodec, a multi-codebook discrete audio codec. It was trained on 7 languages and ~200k hours of audio (mostly tested in English and German).

I wanted to make a post to highlight two choices that worked well for us and seem less common in current TTS models:

Read the original at Machine Learning