Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions
Shuhao Gu, Bojie Hu, Yang Feng
摘要
This paper considers continual learning of large-scale pretrained neural machine translation model without accessing the previous training data or introducing model separation. We argue that the widely used regularization-based methods, which perform multi-objective learning with an auxiliary loss, suffer from the misestimate problem and cannot always achieve a good balance between the previous and new tasks. To solve the problem, we propose a two-stage training method based on the local features of the real loss. We first search low forgetting risk regions, where the model can retain the performance on the previous task as the parameters are updated, to avoid the catastrophic forgetting problem. Then we can continually train the model within this region only with the new training data to fit the new task. Specifically, we propose two methods to search the low forgetting risk regions, which are based on the curvature of loss and the impacts of the parameters on the model output, respectively. We conduct experiments on domain adaptation and more challenging language adaptation tasks, and the experimental results show that our method can achieve significant improvements compared with several strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu 等ACL 2024 · 被引用 7 次
- Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Junpeng Liu, Maosong Sun 等EMNLP 2023 · 被引用 4 次
- Continual Learning for Multilingual Neural Machine Translation via Dual Importance-based Model DivisionJunpeng Liu, Kaiyu Huang, Hao Yu, Jiuyi Li 等EMNLP 2023 · 被引用 4 次
- TL-CL: Task And Language Incremental Continual LearningShrey Satapara, P. K. SrijithEMNLP 2024 · 被引用 3 次
它引用的顶会 Paper4
- Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine TranslationMozhdeh Gheini, Xiang Ren, Jonathan MayEMNLP 2021 · 被引用 133 次
- Boosting Neural Machine Translation with Similar TranslationsJitao Xu, Josep Maria Crego, Jean SenellartACL 2020 · 被引用 59 次
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 被引用 13 次
- Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary 等ACL 2021
相关 Paper
- Compositional Language Continual LearningYuanpeng Li, Liang Zhao, Kenneth Church, Mohamed ElhoseinyICLR 2020 · 被引用 40 次
- Domain adapted machine translation: What does catastrophic forgetting forget and why?Danielle Saunders, Steve DeNeefeEMNLP 2024 · 被引用 1 次
- Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language ModelsYuehao Liu, Shanyan Guan, Weijia Zhang, Xuanming Shang 等CVPR 2026
- Gradient Regularized Contrastive Learning for Continual Domain AdaptationShixiang Tang, Peng Su, Dapeng Chen, Wanli OuyangAAAI 2021 · 被引用 72 次
- LCA: Local Classifier Alignment for Continual LearningTung Tran, Danilo Vasconcellos Vargas, Khoat ThanICLR 2026
