Condensing Multilingual Knowledge with Lightweight Language-Specific Modules
Haoran Xu, Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray
Abstract
Incorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage. We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue. LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation. Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency. Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution 1 We release our code at: https://github.com/fe1ixxu/ LMS_FD . 2 Each pass through the model utilizes only the corresponding language-specific component. The additional computational cost may only come from communication among devices (such as ALLToALL) or gate routing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e0dd7bc-a944-4dbc-a0db-53ed80e6861aCited by top-tier papers1
Ask how each one uses itBuilds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
Related papers
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationTianyu Dong, Bo Li, Jinsong Liu, Shaolin Zhu et al.ACL 2025
- Sparse MoE with Language Guided Routing for Multilingual Machine TranslationXinyu Zhao, Xuxi Chen, Yu Cheng, Tianlong ChenICLR 2024 · 19 citations
- MoLoRA: Boosting LLM-based End-to-end Speech Translation with Mixture of Low-rank ExpertsHao Zhang, Yaqi Chen, Nianwen Si, XuKui Yang et al.AAAI 2026
- A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She et al.ACL 2026
- Unifying the Convergences in Multilingual Neural Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Bing QinEMNLP 2022 · 6 citations
