Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation
Biao Zhang, Ankur Bapna, Rico Sennrich, Orhan Firat
摘要
Using a mix of shared and language-specific (LS) parameters has shown promise in multilingual neural machine translation (MNMT), but the question of when and where LS capacity matters most is still under-studied. We offer such a study by proposing conditional language-specific routing (CLSR). CLSR employs hard binary gates conditioned on token representations to dynamically select LS or shared paths. By manipulating these gates, it can schedule LS capacity across sub-layers in MNMT subject to the guidance of translation signals and budget constraints. Moreover, CLSR can easily scale up to massively multilingual settings. Experiments with Transformer on OPUS-100 and WMT datasets show that: 1) MNMT is sensitive to both the amount and the position of LS modeling: distributing 10%-30% LS computation to the top and/or bottom encoder/decoder layers delivers the best performance; and 2) one-to-many translation benefits more from CLSR compared to many-to-one translation, particularly with unbalanced training data. Our study further verifies the trade-off between the shared capacity and LS capacity for multilingual translation. We corroborate our analysis by confirming the soundness of our findings as foundation of our improved multilingual Transformers. Source code and models are available at https://github.com/bzhangGo/zero/tree/iclr2021_clsr. Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics One-sentence Summary: We investigate and improve parameter-sharing strategies in multilingual Transformers by utilizing conditional computation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper34
- Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEsJinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang 等NeurIPS 2022 · 被引用 95 次
- MLSLT: Towards Multilingual Sign Language TranslationAoxiong Yin, Zhou Zhao, Weike Jin, Meng Zhang 等CVPR 2022 · 被引用 46 次
- Examining Scaling and Transfer of Language Model Architectures for Machine TranslationBiao Zhang, Behrooz Ghorbani, Ankur Bapna, Yong Cheng 等ICML 2022 · 被引用 30 次
- Robust Optimization for Multilingual Translation with Imbalanced DataXian Li, Hongyu GongNeurIPS 2021 · 被引用 24 次
- Parameter Differentiation Based Multilingual Neural Machine TranslationQian Wang, Jiajun ZhangAAAI 2022 · 被引用 21 次
相关 Paper
- LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine TranslationShaolin Zhu, Leiyu Pan, Bo Li, Deyi XiongACL 2024
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu 等ICLR 2026 · 被引用 34 次
- Learning Language-Specific Layers for Multilingual Machine TranslationTelmo Pires, Robin M. Schmidt, Yi-Hsiu Liao, Stephan PeitzACL 2023 · 被引用 8 次
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 被引用 2 次
- MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksShangjie Li, Xiangpeng Wei, Shaolin Zhu, Jun Xie 等EMNLP 2023 · 被引用 4 次
