natural language processing

Six subtitle models, six languages: metrics only told part of the story.

In our recent benchmark, we evaluated TranslateGemma against five other leading LLMs for subtitle translation across six languages: Spanish, Japanese, Korean, Thai, Chinese Simplified, and Chinese Traditional.

3 min readMachine Learning
Six subtitle models, six languages: metrics only told part of the story.
We benchmarked TranslateGemma against 5 other LLMs on subtitle translation across 6 languages. At first glance the numbers told a clean story, but then human QA added a chapter. [D]

The numbers told a clean story, and the story was wrong. TranslateGemma-12b topped our six-model benchmark across all six languages, and the gap was not subtle. But when our linguists reviewed the Traditional Chinese output, the model was quietly producing Simplified Chinese for both zh-CN and zh-TW tags. MetricX-24 and COMETKiwi scored that output identically and highly. No flag, no warning, no hint that anything was off. The metrics did their job, and the job was not the one we needed done.

This is the part of the story that matters for anyone building a translation pipeline. TranslateGemma's fine-tuning corpus is heavily skewed toward Simplified Chinese, and the model has learned to honor that bias over the explicit locale tag. This is a confirmed, publicly documented issue affecting all model sizes, 4B, 12B, and 27B. Upgrading to a larger parameter count will not help because the root cause is training data composition, not capacity. The workaround exists, OpenCC s2twp post-processing, but the deeper problem is that standard QE metrics will look fine the entire time. You can build an automated validation layer, watch it pass every checkpoint, and still ship the wrong script to your users.

The metric-model affinity concern adds another layer. MetricX-24 is a Google metric, TranslateGemma is a Google model, and the gap between them narrows noticeably when you look at COMETKiwi from Unbabel. The direction of the result holds either way, but the size of the lead is likely inflated. Claude's collapse in Japanese tells a similar story from the opposite angle: its COMETKiwi was decent at 0.79, but its MetricX score of 3.90 was the worst across all languages. The output sounded natural and drifted from the source meaning. Fluency and fidelity are not the same thing, and if your quality gate only measures one, you are flying blind in the other direction.

The practical takeaway is not that TranslateGemma is a bad model or that QE metrics are useless. It is that automated evaluation is a necessary filter, not a final arbiter. The TQI composite we developed, COMETKiwi multiplied by an exponential decay of MetricX, gave us a useful ranking signal. It did not catch a script-level failure that a human reviewer spotted in minutes. If you are deploying these models in production, budget for human QA on the edge cases, especially for language pairs where training data bias is a known risk. The metrics will tell you which model to use. Only a human will tell you whether the output is actually what your users need.

From Machine Learning

We evaluated six models on English subtitle translation into Spanish, Japanese, Korean, Thai, Chinese Simplified, and Chinese Traditional - 167 segments per language pair, scored with two reference-free QE metrics.

We used MetricX-24 (lower = better) and COMETKiwi (higher = better) - both reference-free QE metrics. We also developed a combined score:

Read the original at Machine Learning