Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation
Xinyi Wang, Sebastian Ruder, Graham Neubig
Abstract
The performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text present in a target language. Thus, the majority of the world's languages cannot benefit from recent progress in NLP as they have no or limited textual data. To expand possibilities of using NLP technology in these under-represented languages, we systematically study strategies that relax the reliance on conventional language resources through the use of bilingual lexicons, an alternative resource with much better language coverage. We analyze different strategies to synthesize textual or labeled data using lexicons, and how this data can be combined with monolingual or parallel text when available. For 19 under-represented languages across 3 tasks, our methods lead to consistent improvements of up to 5 and 15 points with and without extra monolingual text respectively. Overall, our study highlights how NLP methods can be adapted to thousands more languages that are under-served by current technology. 1 1 Code and data are available at: https: //github.com/cindyxinyiwang/ expand-via-lexicon-based-adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a01b7a6a-068d-4d31-a55b-308cf5eb2cbdCited by top-tier papers13
- A Benchmark for Learning to Translate a New Language from One Grammar BookGarrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky et al.ICLR 2024 · 97 citations
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
- Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchACL 2025 · 17 citations
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 16 citations
- Glot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesAyyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini et al.ACL 2023 · 14 citations
Builds on10
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
Related papers
- GATITOS: Using a New Multilingual Lexicon for Low-resource Machine TranslationAlexander Jones, Isaac Caswell, Orhan Firat, Ishank SaxenaEMNLP 2023 · 3 citations
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 12 citations
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia et al.ACL 2024 · 1 citation
- Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary et al.ACL 2021
