Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation
Chenze Shao, Yang Feng
摘要
Neural networks tend to gradually forget the previously learned knowledge when learning multiple tasks sequentially from dynamic data distributions. This problem is called catastrophic forgetting, which is a fundamental challenge in the continual learning of neural networks. In this work, we observe that catastrophic forgetting not only occurs in continual learning but also affects the traditional static training. Neural networks, especially neural machine translation models, suffer from catastrophic forgetting even if they learn from a static training set. To be specific, the final model pays imbalanced attention to training samples, where recently exposed samples attract more attention than earlier samples. The underlying cause is that training samples do not get balanced training in each model update, so we name this problem imbalanced training. To alleviate this problem, we propose Complementary Online Knowledge Distillation (COKD), which uses dynamically updated teacher models trained on specific data orders to iteratively provide complementary knowledge to the student model. Experimental results on multiple machine translation tasks show that our method successfully alleviates the problem of imbalanced training and achieves substantial improvements over strong baseline systems. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 被引用 304 次
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu 等ASE 2023 · 被引用 91 次
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- Order Matters in the Presence of Dataset Imbalance for Multilingual LearningDami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer 等NeurIPS 2023 · 被引用 9 次
- MultiSFL: Towards Accurate Split Federated Learning via Multi-Model Aggregation and Knowledge ReplayZeke Xia, Ming Hu, Dengke Yan, Ruixuan Liu 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper6
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
- Peer Collaborative Learning for Online Knowledge DistillationGuile Wu, Shaogang GongAAAI 2021 · 被引用 150 次
- Finding Sparse Structures for Domain Specific Neural Machine TranslationJianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu 等AAAI 2021 · 被引用 33 次
- Vocabulary Learning via Optimal Transport for Neural Machine TranslationJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng 等ACL 2021
- Importance-based Neuron Allocation for Multilingual Neural Machine TranslationWanying Xie, Yang Feng, Shuhao Gu, Dong YuACL 2021
相关 Paper
- Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationYunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu 等ACL 2024 · 被引用 7 次
- Gradient Reweighting: Towards Imbalanced Class-Incremental LearningJiangpeng HeCVPR 2024
- Static-Dynamic Co-teaching for Class-Incremental 3D Object DetectionNa Zhao, Gim Hee LeeAAAI 2022 · 被引用 26 次
- Learning without Prejudices: Continual Unbiased Learning via Benign and Malignant ForgettingMyeongho Jeon, Hyoje Lee, Yedarm Seong, Myungjoo KangICLR 2023
- Bridging Non Co-occurrence with Unlabeled In-the-wild Data for Incremental Object DetectionNa Dong, Yongqiang Zhang, Mingli Ding, Gim Hee LeeNeurIPS 2021 · 被引用 33 次
