SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani
Abstract
Evaluating machine translation (MT) quality for under-resourced African languages remains a significant challenge, as existing metrics often suffer from limited language coverage and poor performance in low-resource settings. While recent efforts, such as AfriCOMET, have addressed some of the issues, they are still constrained by small evaluation sets, a lack of publicly available training data tailored to African languages, and inconsistent performance in extremely low-resource scenarios. In this work, we introduce SSA-MTE, a large-scale humanannotated MT evaluation (MTE) dataset covering 14 African language pairs from the News domain, with over 73,000 sentence-level annotations from a diverse set of MT systems. Based on this data, we develop SSA-COMET and SSA-COMET-QE, improved reference-based and reference-free evaluation metrics. We also benchmark prompting-based approaches using state-of-the-art LLMs like GPT-4o, Claude-3.7 and Gemini 2.5 Pro . Our experimental results show that SSA-COMET models significantly outperform AfriCOMET and are competitive with the strongest LLM (Gemini 2.5 Pro) evaluated in our study, particularly on low-resource languages such as Twi, Luo, and Yorùbá. All resources are released under open licenses to support future research. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1eef1b3f-2aad-4009-a9d6-abd62c0250cfCited by top-tier papers1
Ask how each one uses itBuilds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan et al.ACL 2023 · 8 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- UniTE: Unified Translation EvaluationYu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang et al.ACL 2022
Related papers
- AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African LanguagesMachel Reid, Junjie Hu, Graham Neubig, Yutaka MatsuoEMNLP 2021 · 12 citations
- What do Large Language Models Need for Machine Translation Evaluation?Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia et al.EMNLP 2024 · 4 citations
- AFRIDOC-MT: Document-level MT Corpus for African LanguagesJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet et al.EMNLP 2025
- INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African LanguagesHao Yu, Jesujoba Oluwadara Alabi, Andiswa Bukula, Jian Yun Zhuang et al.ACL 2025 · 8 citations
- Error Analysis of Multilingual Language Models in Machine Translation: A Case Study of English-Amharic TranslationHizkiel Alemayehu, Hamada M. Zahera, Axel-Cyrille Ngonga NgomoEMNLP 2024 · 1 citation
