XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, Melvin Johnson
Abstract
Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We release the benchmark 1 to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers139
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- AraT5: Text-to-Text Transformers for Arabic Language GenerationEl Moatez Billah Nagoudi, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2022 · 175 citations
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferIlias Chalkidis, Manos Fergadiotis, Ion AndroutsopoulosEMNLP 2021 · 78 citations
- Dual-Objective Fine-Tuning of BERT for Entity MatchingRalph Peeters, Christian BizerVLDB 2021 · 71 citations
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy et al.ICML 2022 · 71 citations
Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine TranslationAditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari et al.AAAI 2020 · 74 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant et al.EMNLP 2021 · 10 citations
- Enhancing Cross-lingual Transfer by Manifold MixupHuiyun Yang, Huadong Chen, Hao Zhou, Lei LiICLR 2022 · 49 citations
- MCL-NER: Cross-Lingual Named Entity Recognition via Multi-View Contrastive LearningYing Mo, Jian Yang, Jiahao Liu, Qifan Wang et al.AAAI 2024 · 42 citations
- P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMsYidan Zhang, Yu Wan, Boyi Deng, Baosong Yang et al.EMNLP 2025
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 4 citations
