MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
Shun Wang, Ge Zhang, Han Wu, Tyler Loakman, Wenhao Huang, Chenghua Lin
Abstract
Machine Translation (MT) has developed rapidly since the release of Large Language Models and current MT evaluation is performed through comparison with reference human translations or by predicting quality scores from human-labeled data. However, these mainstream evaluation methods mainly focus on fluency and factual reliability, whilst paying little attention to figurative quality. In this paper, we investigate the figurative quality of MT and propose a set of human evaluation metrics focused on the translation of figurative language. We additionally present a multilingual parallel metaphor corpus generated by postediting. Our evaluation protocol is designed to estimate four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality. In doing so, we observe that translations of figurative expressions display different traits from literal ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1912417d-7827-4e42-8ddc-7760354cb726Cited by top-tier papers1
Ask how each one uses itBuilds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile GenerationTuhin Chakrabarty, Smaranda Muresan, Nanyun PengEMNLP 2020 · 46 citations
- Should a Chatbot be Sarcastic? Understanding User Preferences Towards Sarcasm GenerationSilviu Vlad Oprea, Steven R. Wilson, Walid MagdyACL 2022 · 12 citations
- Learning Confidence for Transformer-based Neural Machine TranslationYu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu et al.ACL 2022 · 9 citations
Related papers
- Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation EvaluationYanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang et al.ACL 2026 · 3 citations
- FiRE: Fine-grained Ranking Evaluation for Machine TranslationWenyang Gao, Yinghao Yang, Xi Jin, Jing Li et al.ICML 2026
- Don't Go Far Off: An Empirical Study on Neural Poetry TranslationTuhin Chakrabarty, Arkadiy Saakyan, Smaranda MuresanEMNLP 2021 · 8 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 13 citations
