Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model Division
Junpeng Liu, Kaiyu Huang, Hao Yu, Jiuyi Li, Jinsong Su, Degen Huang
摘要
A persistent goal of multilingual neural machine translation (MNMT) is to continually adapt the model to support new language pairs or improve some current language pairs without accessing the previous training data. To achieve this, the existing methods primarily focus on preventing catastrophic forgetting by making compromises between the original and new language pairs, leading to sub-optimal performance on both translation tasks. To mitigate this problem, we propose a dual importancebased model division method to divide the model parameters into two parts and separately model the translation of the original and new tasks. Specifically, we first remove the parameters that are negligible to the original tasks but essential to the new tasks to obtain a pruned model, which is responsible for the original translation tasks. Then we expand the pruned model with external parameters and fine-tune the newly added parameters with new training data. The whole fine-tuned model will be used for the new translation tasks. Experimental results show that our method can efficiently adapt the original model to various new translation tasks while retaining the performance of the original tasks. Further analyses demonstrate that our method consistently outperforms several strong baselines under different incremental translation scenarios. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive ProjectionSaleh Momeni, Changnan Xiao, Bing LiuNeurIPS 2025 · 被引用 8 次
- A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language ModelsZhihao Wang, Shiyu Liu, Jianheng Huang, Wang Zheng 等EMNLP 2024
它引用的顶会 Paper11
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- Learning Sparse Sharing Architectures for Multiple TasksTianxiang Sun, Yunfan Shao, Xiaonan Li, Pengfei Liu 等AAAI 2020 · 被引用 155 次
- Continual Learning in Task-Oriented Dialogue SystemsAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon 等EMNLP 2021 · 被引用 68 次
相关 Paper
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- Importance-based Neuron Allocation for Multilingual Neural Machine TranslationWanying Xie, Yang Feng, Shuhao Gu, Dong YuACL 2021
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Junpeng Liu, Maosong Sun 等EMNLP 2023 · 被引用 4 次
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 被引用 2 次
- Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsShuhao Gu, Bojie Hu, Yang FengEMNLP 2022 · 被引用 11 次
