Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions
Shuhao Gu, Bojie Hu, Yang Feng
Abstract
This paper considers continual learning of large-scale pretrained neural machine translation model without accessing the previous training data or introducing model separation. We argue that the widely used regularization-based methods, which perform multi-objective learning with an auxiliary loss, suffer from the misestimate problem and cannot always achieve a good balance between the previous and new tasks. To solve the problem, we propose a two-stage training method based on the local features of the real loss. We first search low forgetting risk regions, where the model can retain the performance on the previous task as the parameters are updated, to avoid the catastrophic forgetting problem. Then we can continually train the model within this region only with the new training data to fit the new task. Specifically, we propose two methods to search the low forgetting risk regions, which are based on the curvature of loss and the impacts of the parameters on the model output, respectively. We conduct experiments on domain adaptation and more challenging language adaptation tasks, and the experimental results show that our method can achieve significant improvements compared with several strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1296e7d9-7912-4f30-8865-bc63475fdda7Cited by top-tier papers6
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu et al.ACL 2024 · 7 citations
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Junpeng Liu, Maosong Sun et al.EMNLP 2023 · 4 citations
- Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionJunpeng Liu, Kaiyu Huang, Hao Yu, Jiuyi Li et al.EMNLP 2023 · 4 citations
- TL-CL: Task And Language Incremental Continual LearningShrey Satapara, P. K. SrijithEMNLP 2024 · 3 citations
Builds on4
- Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine TranslationMozhdeh Gheini, Xiang Ren, Jonathan MayEMNLP 2021 · 133 citations
- Boosting Neural Machine Translation with Similar TranslationsJitao Xu, Josep Maria Crego, Jean SenellartACL 2020 · 59 citations
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 13 citations
- Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary et al.ACL 2021
Related papers
- Compositional Language Continual LearningYuanpeng Li, Liang Zhao, Kenneth Church, Mohamed ElhoseinyICLR 2020 · 40 citations
- Domain adapted machine translation: What does catastrophic forgetting forget and why?Danielle Saunders, Steve DeNeefeEMNLP 2024 · 1 citation
- Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language ModelsYuehao Liu, Shanyan Guan, Weijia Zhang, Xuanming Shang et al.CVPR 2026
- Gradient Regularized Contrastive Learning for Continual Domain AdaptationShixiang Tang, Peng Su, Dapeng Chen, Wanli OuyangAAAI 2021 · 72 citations
- LCA: Local Classifier Alignment for Continual LearningTung Tran, Danilo Vasconcellos Vargas, Khoat ThanICLR 2026
