Discovering Language-neutral Sub-networks in Multilingual Language Models
Negar Foroutan, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut, Karl Aberer
Abstract
Multilingual pre-trained language models transfer remarkably well on cross-lingual downstream tasks. However, the extent to which they learn language-neutral representations (i.e., shared representations that encode similar phenomena across languages), and the effect of such representations on cross-lingual transfer performance, remain open questions. In this work, we conceptualize language neutrality of multilingual models as a function of the overlap between language-encoding subnetworks of these models. We employ the lottery ticket hypothesis to discover sub-networks that are individually optimized for various languages and tasks. Our evaluation across three distinct tasks and eleven typologically-diverse languages demonstrates that sub-networks for different languages are topologically similar (i.e., language-neutral), making them effective initializations for cross-lingual transfer with limited performance degradation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9a677df-429c-4d12-87da-3483f73e437dCited by top-tier papers5
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
- Improving the Cross-Lingual Generalisation in Visual Question AnsweringFarhad Nooralahzadeh, Rico SennrichAAAI 2023 · 8 citations
- How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuningRochelle Choenni, Dan Garrette, Ekaterina ShutovaEMNLP 2023 · 2 citations
- Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron AnalysisZeping Yu, Sophia AnaniadouEMNLP 2024 · 1 citation
- Efficient Unseen Language Adaptation for Multilingual Pre-Trained Language ModelsPo-Heng Chen, Yun-Nung ChenEMNLP 2024
Builds on13
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Training Neural Networks with Fixed Sparse MasksYi-Lin Sung, Varun Nair, Colin RaffelNeurIPS 2021 · 295 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
Related papers
- Playing Lottery Tickets with Vision and LanguageZhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen et al.AAAI 2022 · 64 citations
- Cross-Lingual Generalization and Compression: From Language-Specific to Shared NeuronsFrederick Riemenschneider, Anette FrankACL 2025 · 3 citations
- The Geometry of Multilingual Language Model RepresentationsTyler A. Chang, Zhuowen Tu, Benjamin K. BergenEMNLP 2022 · 22 citations
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving GeneralizationChen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang et al.ACL 2021
- Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationJaeseong Lee, Dohyeon Lee, Seung-won HwangAAAI 2023 · 1 citation
