Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, Hinrich Schütze
摘要
The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An important part of this effort is to collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m. We evaluate Glot500-m on five diverse tasks across these languages. We observe large improvements for both high-resource and low-resource languages compared to an XLM-R baseline. Our analysis shows that no single factor explains the quality of multilingual LLM representations. Rather, a combination of factors determines quality including corpus size, script, "help" from related languages and the total capacity of the model. Our work addresses an important goal of NLP research: we should not limit NLP to a small fraction of the world's languages and instead strive to support as many languages as possible to bring the benefits of NLP technology to all languages and cultures. Code, data and models are available at https://github.com/cisnlp/Glot500 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Constrained Decoding for Cross-lingual Label ProjectionDuong Minh Le, Yang Chen, Alan Ritter, Wei XuICLR 2024 · 被引用 14 次
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 被引用 12 次
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 被引用 3 次
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov 等ACL 2025 · 被引用 1 次
它引用的顶会 Paper6
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu 等AAAI 2020 · 被引用 94 次
- Expanding Pretrained Models to Thousands More Languages via Lexicon-based AdaptationXinyi Wang, Sebastian Ruder, Graham NeubigACL 2022 · 被引用 73 次
- Masking as an Efficient Alternative to Finetuning for Pretrained Language ModelsMengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi 等EMNLP 2020 · 被引用 61 次
- A Large-Scale Study of Machine Translation in Turkic LanguagesJamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev 等EMNLP 2021 · 被引用 19 次
相关 Paper
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language ModelsDavis Liang, Hila Gonen, Yuning Mao, Rui Hou 等EMNLP 2023 · 被引用 29 次
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang 等ICLR 2025
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao 等ACL 2026 · 被引用 11 次
- An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo 等ACL 2025
