Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
Klaudia Thellmann, Bernhard Stadler, Michael Färber, Jens Lehmann
Abstract
Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these benchmarks remain underexplored, raising concerns about the reliability and comparability of multilingual evaluation. We address two practical gaps: (i) how well automatic MQM-style error spans from LLM judges and a span-aware QE baseline (xCOMET-XXL) match expert human span annotations on benchmark translations, and (ii) how strongly translation errors (as opposed to source-side issues in the English original) explain accuracy drops on translated benchmarks. We find that span agreement is non-trivial on naturally occurring benchmark translations, and that target-side translation errors are consistently associated with measurable, percentage-point drops in translated accuracy even after controlling for English correctness and source-side anomalies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f82c02f-4816-4aed-bf9c-4ff5a69e8152Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes et al.ACL 2025
- Enhancing Human Evaluation in Machine Translation with Comparative JudgementYixiao Song, Parker Riley, Daniel Deutsch, Markus FreitagACL 2025 · 2 citations
- Lost in Literalism: How Supervised Training Shapes Translationese in LLMsYafu Li, Ronghao Zhang, Zhilin Wang, Huajian Zhang et al.ACL 2025 · 12 citations
- Revisiting Machine Translation for Cross-lingual ClassificationMikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan et al.EMNLP 2023 · 10 citations
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question AnsweringKeon-Woo Roh, Yeong-Joon Ju, Seong-Whan LeeEMNLP 2025
