What do Large Language Models Need for Machine Translation Evaluation?
Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Frédéric Blain
摘要
Leveraging large language models (LLMs) for various natural language processing tasks has led to superlative claims about their performance. For the evaluation of machine translation (MT), existing research shows that LLMs are able to achieve results comparable to finetuned multilingual pre-trained language models. In this paper, we explore what translation information, such as the source, reference, translation errors and annotation guidelines, is needed for LLMs to evaluate MT quality. In addition, we investigate prompting techniques such as zero-shot, Chain of Thought (CoT) and few-shot prompting for eight language pairs covering high-, medium-and lowresource languages, leveraging varying LLM variants. Our findings indicate the importance of reference translations for an LLM-based evaluation. While larger models do not necessarily fare better, they tend to benefit more from CoT prompting, than smaller models. We also observe that LLMs do not always provide a numerical score when generating evaluations, which poses a question on their reliability for the task. Our work presents a comprehensive analysis for resource-constrained and trainingless LLM-based evaluation of machine translation. We release the accrued prompt templates, code and data publicly for reproducibility 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward ModelingShaomu Tan, Christof MonzEMNLP 2025
- MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song TranslationWoohyun Cho, Youngmin Kim, Sunghyun Lee, Youngjae YuEMNLP 2025
- DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationZhibo Man, Yuanmeng Chen, Yujie Zhang, Jinan XuEMNLP 2025
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- OpenChat: Advancing Open-source Language Models with Mixed-Quality DataGuan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li 等ICLR 2024 · 被引用 328 次
相关 Paper
- PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization EvaluationChristoph Leiter, Steffen EgerEMNLP 2024 · 被引用 5 次
- Prompting PaLM for Translation: Assessing Strategies and PerformanceDavid Vilar, Markus Freitag, Colin Cherry, Jiaming Luo 等ACL 2023 · 被引用 70 次
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 被引用 402 次
- On Bilingual Lexicon Induction with Large Language ModelsYaoyiran Li, Anna Korhonen, Ivan VulicEMNLP 2023 · 被引用 2 次
- SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry 等EMNLP 2025
