Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Wenbo Yu, Dawei Song, Jie Tang, Yuhang Guo
Abstract
Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc. In these scenarios, translations often require handling non-literal expressions, leading to the inaccuracy of MT metrics. To systematically investigate the reliability of MT metrics, we first curate a meta-evaluation dataset focused on non-literal translations, namely MENT. MENT encompasses four non-literal translation domains and features source sentences paired with translations from diverse MT systems, with 7,530 human-annotated scores on translation quality. Experimental results reveal the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge, particularly the knowledge cutoff and score inconsistency problem. To mitigate these limitations, we propose RATE, a novel agentic translation evaluation framework, centered by a reflective Core Agent that dynamically invokes specialized sub-agents. Experimental results indicate the efficacy of RATE, achieving an improvement of at least 3.2 points in combined system- and segment-level correlation with human judgments compared with current methods. Further experiments demonstrate the robustness of RATE to general-domain MT evaluation. Code and dataset are available at: https://github.com/BITHLP/RATE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 122 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsRicardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves et al.ACL 2026 · 34 citations
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 14 citations
Related papers
- M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation EvaluationZhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren et al.ACL 2025
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman et al.EMNLP 2024 · 3 citations
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceAndong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai et al.EMNLP 2025
- Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li et al.ACL 2026 · 2 citations
- Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance BoostRunzhe Zhan, Zhihong Huang, Xinyi Yang, Lidia S. Chao et al.NeurIPS 2025 · 6 citations
