Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, Hinrich Schütze
Abstract
The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An important part of this effort is to collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m. We evaluate Glot500-m on five diverse tasks across these languages. We observe large improvements for both high-resource and low-resource languages compared to an XLM-R baseline. Our analysis shows that no single factor explains the quality of multilingual LLM representations. Rather, a combination of factors determines quality including corpus size, script, "help" from related languages and the total capacity of the model. Our work addresses an important goal of NLP research: we should not limit NLP to a small fraction of the world's languages and instead strive to support as many languages as possible to bring the benefits of NLP technology to all languages and cultures. Code, data and models are available at https://github.com/cisnlp/Glot500 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cd3e582-8768-4638-81a7-4ff6e13849acCited by top-tier papers14
- Constrained Decoding for Cross-lingual Label ProjectionDuong Minh Le, Yang Chen, Alan Ritter, Wei XuICLR 2024 · 14 citations
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 12 citations
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 3 citations
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov et al.ACL 2025 · 1 citation
Builds on6
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu et al.AAAI 2020 · 94 citations
- Expanding Pretrained Models to Thousands More Languages via Lexicon-based AdaptationXinyi Wang, Sebastian Ruder, Graham NeubigACL 2022 · 73 citations
- Masking as an Efficient Alternative to Finetuning for Pretrained Language ModelsMengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi et al.EMNLP 2020 · 61 citations
- A Large-Scale Study of Machine Translation in Turkic LanguagesJamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev et al.EMNLP 2021 · 19 citations
Related papers
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsDavis Liang, Hila Gonen, Yuning Mao, Rui Hou et al.EMNLP 2023 · 29 citations
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang et al.ICLR 2025
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao et al.ACL 2026 · 11 citations
- An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo et al.ACL 2025
