Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation
Chenze Shao, Yang Feng
Abstract
Neural networks tend to gradually forget the previously learned knowledge when learning multiple tasks sequentially from dynamic data distributions. This problem is called catastrophic forgetting, which is a fundamental challenge in the continual learning of neural networks. In this work, we observe that catastrophic forgetting not only occurs in continual learning but also affects the traditional static training. Neural networks, especially neural machine translation models, suffer from catastrophic forgetting even if they learn from a static training set. To be specific, the final model pays imbalanced attention to training samples, where recently exposed samples attract more attention than earlier samples. The underlying cause is that training samples do not get balanced training in each model update, so we name this problem imbalanced training. To alleviate this problem, we propose Complementary Online Knowledge Distillation (COKD), which uses dynamically updated teacher models trained on specific data orders to iteratively provide complementary knowledge to the student model. Experimental results on multiple machine translation tasks show that our method successfully alleviates the problem of imbalanced training and achieves substantial improvements over strong baseline systems. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ec8e906-e9c2-465f-aa39-30788a2442ccCited by top-tier papers12
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 304 citations
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
- Order Matters in the Presence of Dataset Imbalance for Multilingual LearningDami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer et al.NeurIPS 2023 · 9 citations
- MultiSFL: Towards Accurate Split Federated Learning via Multi-Model Aggregation and Knowledge ReplayZeke Xia, Ming Hu, Dengke Yan, Ruixuan Liu et al.AAAI 2025 · 8 citations
Builds on6
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Peer Collaborative Learning for Online Knowledge DistillationGuile Wu, Shaogang GongAAAI 2021 · 150 citations
- Finding Sparse Structures for Domain Specific Neural Machine TranslationJianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu et al.AAAI 2021 · 33 citations
- Vocabulary Learning via Optimal Transport for Neural Machine TranslationJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng et al.ACL 2021
- Importance-based Neuron Allocation for Multilingual Neural Machine TranslationWanying Xie, Yang Feng, Shuhao Gu, Dong YuACL 2021
Related papers
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu et al.ACL 2024 · 7 citations
- Gradient Reweighting: Towards Imbalanced Class-Incremental LearningJiangpeng HeCVPR 2024
- Static-Dynamic Co-teaching for Class-Incremental 3D Object DetectionNa Zhao, Gim Hee LeeAAAI 2022 · 26 citations
- Learning without Prejudices: Continual Unbiased Learning via Benign and Malignant ForgettingMyeongho Jeon, Hyoje Lee, Yedarm Seong, Myungjoo KangICLR 2023
- Bridging Non Co-occurrence with Unlabeled In-the-wild Data for Incremental Object DetectionNa Dong, Yongqiang Zhang, Mingli Ding, Gim Hee LeeNeurIPS 2021 · 33 citations
