IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian Languages
Ananya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre
Abstract
The rapid growth of machine translation (MT) systems necessitates meta-evaluations of evaluation metrics to enable selection of those that best reflect MT quality. Unfortunately, most meta-evaluation studies focus on European languages, the observations for which may not always apply to other languages. Indian languages, having over a billion speakers, are linguistically different from them, and to date, there are no such systematic studies focused solely on English to Indian language MT. This paper fills this gap through a Multidimensional Quality Metric (MQM) dataset consisting of 7000 fine-grained annotations, spanning 5 Indian languages and 7 MT systems. We evaluate 16 metrics and show that, pre-trained metrics like COMET have the highest correlations with annotator scores as opposed to n-gram metrics like BLEU. We further leverage our MQM annotations to develop an Indic-COMET metric and show that it outperforms COMET counterparts in both human scores correlations and robustness scores in Indian languages. Additionally, we show that the Indic-COMET can outperform COMET on some unseen Indian languages. We hope that our dataset and analysis will facilitate further research in Indic MT evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary et al.ACL 2025 · 5 citations
- Finding Blind Spots in Evaluator LLMs with Interpretable ChecklistsSumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Sshubam Verma, Mitesh M. KhapraEMNLP 2024 · 3 citations
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
- SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry et al.EMNLP 2025
- Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian LanguagesAmir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri KotoACL 2026
Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language GenerationSamuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio et al.EMNLP 2021 · 85 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
- Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level AggregationPaulo R. Cavalin, Pedro Henrique Domingues, Claudio S. PinhanezAAAI 2025 · 5 citations
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang et al.EMNLP 2025
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
