2 min readfrom Machine Learning

Follow-up on the TranslateGemma subtitle benchmark: human review of segments rated "clean" by MetricX-24 and COMETKiwi [D]

Our take

In a recent follow-up, we conducted a human review of the segments rated "clean" by the MetricX-24 and COMETKiwi benchmarks, initially placing TranslateGemma-12b at the top for all language pairs. While these metrics correlate with human judgment overall, we sought to determine the accuracy of their "clean" designations. Analyzing 21 English subtitle segments, we found that 71% of segments flagged as clean contained errors upon human review. Notably, the Japanese translations showed significant mistranslations, highlighting the need for deeper scrutiny in subtitle evaluation.

When benchmarking AI translation models like TranslateGemma against 5 other LLMs on subtitle translation across 6 languages, initial results can paint a deceptively clean picture. The recent follow-up human review of TranslateGemma-12b's "clean" translations, however, reveals a critical gap in how automated metrics operate in their high-confidence zones. While the TQI index combining MetricX-24 and COMETKiwi consistently ranked TranslateGemma top across languages, professional linguists found errors in 71% of segments deemed pristine by the metrics. This isn't just a minor discrepancy; it highlights a fundamental limitation in relying solely on population-level correlation between metrics and human judgment when assessing individual outputs, especially for high-stakes applications like subtitles where accuracy and fluency directly impact user experience.

The findings expose a significant "metric-blindness" rate. Despite TranslateGemma achieving the highest mean COMETKiwi score (0.863) for Japanese translations, all 15 mistranslations in that language were missed by the automated system. Critically, every single one of the 25 accuracy errors identified by humans fell into the quadrant the metrics labeled "clean." This demonstrates that while these large reference-free QE models excel at broad assessment, they lack the granular sensitivity to catch subtle but significant translation failures, particularly in accuracy categories like mistranslation or omission. The fact that the single auto-flagged segment contained only a style error, while 60 human-flagged segments included accuracy issues, underscores a dangerous misalignment between what the metrics prioritize and what truly matters for translation quality.

This discrepancy should prompt a reevaluation of how we define and measure translation quality in AI systems. The data suggests that automated thresholds, even when confidently applied, may create a false sense of security. For users and developers relying on these tools for critical content, the takeaway is clear: automated QE metrics are valuable tools for initial screening and population analysis, but they cannot replace human oversight for verifying individual output quality, especially in complex domains like technical subtitles. The future of robust AI translation lies not just in pushing model performance higher, but in developing more sophisticated evaluation frameworks that combine the efficiency of automation with the nuanced understanding of human experts. As we explore next-generation translation solutions, how can we design hybrid evaluation systems that bridge this gap between automated scoring and human judgment?

A few weeks ago I shared the results of a benchmark here comparing 6 LLMs on subtitle translation, scored with two reference-free QE metrics - MetricX-24 (~13B mT5-XXL) and COMETKiwi (~10.7B XLM-R-XXL) - combined into a TQI index. Posting a follow-up because we did human review afterwards, and the result is worth discussing.

The original benchmark put TranslateGemma-12b first in every language pair. The natural question: are those high scores accurate, or are the metrics insensitive in their high-confidence zone? These metrics correlate well with human judgment at the population level (that's what they're trained for), but population-level correlation doesn't tell you whether the segments they call "clean" are actually clean.

So we ran the check directly. 21 English subtitle segments from one tutorial video. TranslateGemma's translations into 4 languages (ES, JA, TH, ZH-CN - Korean and Traditional Chinese got dropped). All 84 translations chosen because they passed the dashboard clean-rule (MX < 5 AND CK ≥ 0.70) in all 4 languages simultaneously. Then full MQM annotation by professional linguists - Major/Minor severity, with categories covering accuracy (mistranslation, omission, addition, untranslated), fluency (grammar, punctuation, inconsistency), style, terminology.

Results under the dashboard threshold:

  • Auto-flagged: 1/84
  • Human-flagged: 60/84 any-error, 13/84 Major-only
  • Metric-blindness rate (auto-clean ∩ human-flagged / auto-clean): 59/83 = 71% any-error, 12/83 = 14.5% Major-only
  • All 25 human-found Accuracy-class errors fell in the metric-blind quadrant. Zero overlap with the auto-flagged region (which contained one Style-category Major error).
  • Japanese carries 10 of 15 total mistranslations across the dataset, all metric-blind, despite having the highest mean COMETKiwi (0.863) of the four languages.

Caveat: small n, one model, one content set, so the numbers are directional rather than definitive.

Original thread: [link]
Full benchmark report: in comments.

submitted by /u/ritis88
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article

Related Articles