Importance-based Neuron Allocation for Multilingual Neural Machine Translation
Wanying Xie, Yang Feng, Shuhao Gu, Dong Yu
Abstract
Multilingual neural machine translation with a single model has drawn much attention due to its capability to deal with multiple languages. However, the current multilingual translation paradigm often makes the model tend to preserve the general knowledge, but ignore the language-specific knowledge. Some previous works try to solve this problem by adding various kinds of language-specific modules to the model, but they suffer from the parameter explosion problem and require specialized manual design. To solve these problems, we propose to divide the model neurons into general and language-specific parts based on their importance across languages. The general part is responsible for preserving the general knowledge and participating in the translation of all the languages, while the language-specific part is responsible for preserving the languagespecific knowledge and participating in the translation of some specific languages. Experimental results on several language pairs, covering IWSLT and Europarl corpus datasets, demonstrate the effectiveness and universality of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87abe359-453b-4808-a176-883fdd060d78Cited by top-tier papers14
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 40 citations
- Parameter Differentiation Based Multilingual Neural Machine TranslationQian Wang, Jiajun ZhangAAAI 2022 · 21 citations
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie et al.CVPR 2026 · 14 citations
- DNA: Uncovering Universal Latent Forgery KnowledgeJingtong Dou, Chuancheng Shi, Anqi Yi, Shiming Guo et al.ICML 2026 · 8 citations
- MODEL SHAPLEY: Find Your Ideal Parameter Player via One Gradient BackpropagationChu Xu, Xinke Jiang, Rihong Qiu, Jiaran Gao et al.NeurIPS 2025 · 7 citations
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 97 citations
Related papers
- Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionJunpeng Liu, Kaiyu Huang, Hao Yu, Jiuyi Li et al.EMNLP 2023 · 4 citations
- Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationJunpeng Liu, Kaiyu Huang, Jiuyi Li, Huan Liu et al.EMNLP 2022 · 5 citations
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
