Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
Mohamed Elsayed, A. Rupam Mahmood
Abstract
Deep representation learning methods struggle with continual learning, suffering from both catastrophic forgetting of useful units and loss of plasticity, often due to rigid and unuseful units. While many methods address these two issues separately, only a few currently deal with both simultaneously. In this paper, we introduce Utility-based Perturbed Gradient Descent (UPGD) as a novel approach for the continual learning of representations. UPGD combines gradient updates with perturbations, where it applies smaller modifications to more useful units, protecting them from forgetting, and larger modifications to less useful units, rejuvenating their plasticity. We use a challenging streaming learning setup where continual learning problems have hundreds of non-stationarities and unknown task boundaries. We show that many existing methods suffer from at least one of the issues, predominantly manifested by their decreasing accuracy over tasks. On the other hand, UPGD continues to improve performance and surpasses or is competitive with all methods in all problems. Finally, in extended reinforcement learning experiments with PPO, we show that while Adam exhibits a performance drop after initial learning, UPGD avoids it by addressing both continual learning issues. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise NetworksHojoon Lee, Hyeonseo Cho, Hyunseung Kim, Donghu Kim et al.ICML 2024 · 36 citations
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He et al.NeurIPS 2024 · 27 citations
- Non-Stationary Learning of Neural Networks with Automatic Soft Parameter ResetAlexandre Galashov, Michalis K. Titsias, András György, Clare Lyle et al.NeurIPS 2024 · 18 citations
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learningJiashun Liu, Zihao Wu, Johan S. Obando-Ceron, Pablo Samuel Castro et al.NeurIPS 2025 · 15 citations
- Region-Based Optimization in Continual Learning for Audio Deepfake DetectionYujie Chen, Jiangyan Yi, Cunhang Fan, Jianhua Tao et al.AAAI 2025 · 10 citations
Builds on10
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che et al.EMNLP 2020 · 152 citations
Related papers
- Continual Learning through Retrieval and ImaginationZhen Wang, Liu Liu, Yiqun Duan, Dacheng TaoAAAI 2022 · 45 citations
- Adaptive Plasticity Improvement for Continual LearningYan-Shuo Liang, Wu-Jun LiCVPR 2023
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
- Preserving Linear Separability in Continual Learning by Backward Feature ProjectionQiao Gu, Dongsub Shim, Florian ShkurtiCVPR 2023
- Continual Learning with Global AlignmentXueying Bai, Jinghuan Shang, Yifan Sun, Niranjan BalasubramanianNeurIPS 2024 · 1 citation
