Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation
Biao Zhang, Ankur Bapna, Rico Sennrich, Orhan Firat
Abstract
Using a mix of shared and language-specific (LS) parameters has shown promise in multilingual neural machine translation (MNMT), but the question of when and where LS capacity matters most is still under-studied. We offer such a study by proposing conditional language-specific routing (CLSR). CLSR employs hard binary gates conditioned on token representations to dynamically select LS or shared paths. By manipulating these gates, it can schedule LS capacity across sub-layers in MNMT subject to the guidance of translation signals and budget constraints. Moreover, CLSR can easily scale up to massively multilingual settings. Experiments with Transformer on OPUS-100 and WMT datasets show that: 1) MNMT is sensitive to both the amount and the position of LS modeling: distributing 10%-30% LS computation to the top and/or bottom encoder/decoder layers delivers the best performance; and 2) one-to-many translation benefits more from CLSR compared to many-to-one translation, particularly with unbalanced training data. Our study further verifies the trade-off between the shared capacity and LS capacity for multilingual translation. We corroborate our analysis by confirming the soundness of our findings as foundation of our improved multilingual Transformers. Source code and models are available at https://github.com/bzhangGo/zero/tree/iclr2021_clsr. Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics One-sentence Summary: We investigate and improve parameter-sharing strategies in multilingual Transformers by utilizing conditional computation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ac974a8f-8b4b-45bb-aa7d-b71141b95a50Cited by top-tier papers34
- Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEsJinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang et al.NeurIPS 2022 · 95 citations
- MLSLT: Towards Multilingual Sign Language TranslationAoxiong Yin, Zhou Zhao, Weike Jin, Meng Zhang et al.CVPR 2022 · 46 citations
- Examining Scaling and Transfer of Language Model Architectures for Machine TranslationBiao Zhang, Behrooz Ghorbani, Ankur Bapna, Yong Cheng et al.ICML 2022 · 30 citations
- Robust Optimization for Multilingual Translation with Imbalanced DataXian Li, Hongyu GongNeurIPS 2021 · 24 citations
- Parameter Differentiation Based Multilingual Neural Machine TranslationQian Wang, Jiajun ZhangAAAI 2022 · 21 citations
Related papers
- LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine TranslationShaolin Zhu, Leiyu Pan, Bo Li, Deyi XiongACL 2024
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu et al.ICLR 2026 · 34 citations
- Learning Language-Specific Layers for Multilingual Machine TranslationTelmo Pires, Robin M. Schmidt, Yi-Hsiu Liao, Stephan PeitzACL 2023 · 8 citations
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksShangjie Li, Xiangpeng Wei, Shaolin Zhu, Jun Xie et al.EMNLP 2023 · 4 citations
