When metrics mislead, human review reveals what benchmarks miss

In a recent follow-up, we conducted a human review of the segments rated "clean" by the MetricX-24 and COMETKiwi benchmarks, initially placing TranslateGemma-12b at the top for all language pairs.

3 min readMachine Learning

When benchmarking AI translation models like TranslateGemma against 5 other LLMs on subtitle translation across 6 languages, initial results can paint a deceptively clean picture. The recent follow-up human review of TranslateGemma-12b's "clean" translations, however, reveals a critical gap in how automated metrics operate in their high-confidence zones. While the TQI index combining MetricX-24 and COMETKiwi consistently ranked TranslateGemma top across languages, professional linguists found errors in 71% of segments deemed pristine by the metrics. This isn't just a minor discrepancy; it highlights a fundamental limitation in relying solely on population-level correlation between metrics and human judgment when assessing individual outputs, especially for high-stakes applications like subtitles where accuracy and fluency directly impact user experience.

The findings expose a significant "metric-blindness" rate. Despite TranslateGemma achieving the highest mean COMETKiwi score (0.863) for Japanese translations, all 15 mistranslations in that language were missed by the automated system. Critically, every single one of the 25 accuracy errors identified by humans fell into the quadrant the metrics labeled "clean." This demonstrates that while these large reference-free QE models excel at broad assessment, they lack the granular sensitivity to catch subtle but significant translation failures, particularly in accuracy categories like mistranslation or omission. The fact that the single auto-flagged segment contained only a style error, while 60 human-flagged segments included accuracy issues, underscores a dangerous misalignment between what the metrics prioritize and what truly matters for translation quality.

This discrepancy should prompt a reevaluation of how we define and measure translation quality in AI systems. The data suggests that automated thresholds, even when confidently applied, may create a false sense of security. For users and developers relying on these tools for critical content, the takeaway is clear: automated QE metrics are valuable tools for initial screening and population analysis, but they cannot replace human oversight for verifying individual output quality, especially in complex domains like technical subtitles. The future of robust AI translation lies not just in pushing model performance higher, but in developing more sophisticated evaluation frameworks that combine the efficiency of automation with the nuanced understanding of human experts. As we explore next-generation translation solutions, how can we design hybrid evaluation systems that bridge this gap between automated scoring and human judgment?

From Machine Learning

A few weeks ago I shared the results of a benchmark here comparing 6 LLMs on subtitle translation, scored with two reference-free QE metrics - MetricX-24 (~13B mT5-XXL) and COMETKiwi (~10.7B XLM-R-XXL) - combined into a TQI index. Posting a follow-up because we did human review afterwards, and the result is worth discussing.

Read the original at Machine Learning