Stop punishing perfect transcriptions with smarter STT evaluation scoring.

Introducing gladia-normalization: a solution to a common challenge in STT engine benchmarking.

3 min readMachine Learning

Word error rate has a dirty secret: it punishes perfection. When a speech-to-text engine hears "It's $50 at 3:00PM" and writes exactly that, while a human evaluator expects "it is fifty dollars at 3 pm," the scoring system flags a flawless transcript as a failure. This isn't a niche annoyance. It's a structural flaw in how we measure speech recognition quality, and it distorts every benchmark, every model comparison, and every purchasing decision built on top of those numbers.

The team at Gladia saw this problem clearly. Instead of accepting the noise, they built a normalization library that standardizes both the reference text and the hypothesis before computing WER. Their open-source tool, `gladia-normalization`, runs transcripts through a configurable pipeline that handles the messy realities of language: numbers, currency, time formats, punctuation, and casing. The YAML-defined pipelines are deterministic and version-controllable, which means you know exactly what transformations happened and in what order. That's the kind of transparency that builds trust in an evaluation metric, not just a quick hack to make numbers look better.

What matters most here is the shift in mindset. Normalization isn't cheating the metric; it's cleaning the measuring stick. If you're building a voice assistant, a call center analytics tool, or a meeting transcription service, you need to know whether your engine actually understands speech or just got lucky with formatting. By open-sourcing this under an MIT license, Gladia isn't just solving their own problem. They're giving every developer a shared foundation for fairer, more meaningful STT evaluation. That's a practical step toward comparability across projects and vendors, which benefits everyone who relies on these metrics to make decisions.

The current language support covers English, French, German, Italian, Spanish, and Dutch, with the caveat that non-English presets need refinement. That's an honest admission, and it's the right one. Language is messy, and normalization rules are deeply cultural. Gladia is asking native speakers to contribute and shape the behavior for each language, which is the only way to get this right. If you've been wrestling with WER scores that don't reflect reality, this library is worth a look. Not because it's flashy, but because it addresses a real pain point with a reproducible, auditable solution. The next time someone quotes you a WER number, you'll have a better question to ask: normalized against what, and with which rules?

From Machine Learning

Hey guys! At my company, we've been benchmarking STT engines a lot and kept running into the same issue: WER is penalizing formatting differences that have nothing to do with actual recognition quality. "It's $50" vs "it is fifty dollars", "3:00PM" vs "3 pm". Both perfect transcription, but a terrible error rate.

The fix is normalizing both sides before scoring, but every project we had a different script doing it slightly differently. So we built a proper library and open-sourced it.

Read the original at Machine Learning