Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual Learning
Danruo Deng, Guangyong Chen, Jianye Hao, Qiong Wang, Pheng-Ann Heng
Abstract
The backpropagation networks are notably susceptible to catastrophic forgetting, where networks tend to forget previously learned skills upon learning new ones. To address such the 'sensitivity-stability' dilemma, most previous efforts have been contributed to minimizing the empirical risk with different parameter regularization terms and episodic memory, but rarely exploring the usages of the weight loss landscape. In this paper, we investigate the relationship between the weight loss landscape and sensitivity-stability in the continual learning scenario, based on which, we propose a novel method, Flattening Sharpness for Dynamic Gradient Projection Memory (FS-DGPM). In particular, we introduce a soft weight to represent the importance of each basis representing past tasks in GPM, which can be adaptively learned during the learning process, so that less important bases can be dynamically released to improve the sensitivity of new skill learning. We further introduce Flattening Sharpness (FS) to reduce the generalization gap by explicitly regulating the flatness of the weight loss landscape of all seen tasks. As demonstrated empirically, our proposed method consistently outperforms baselines with the superior ability to learn new skills while alleviating forgetting effectively. 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6b030a8-0e32-429f-987c-6220e395a58aCited by top-tier papers43
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
- Loss Decoupling for Task-Agnostic Continual LearningYan-Shuo Liang, Wu-Jun LiNeurIPS 2023 · 63 citations
- A Unified Approach to Domain Incremental Learning with Memory: Theory and AlgorithmHaizhou Shi, Hao WangNeurIPS 2023 · 60 citations
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu et al.NeurIPS 2024 · 48 citations
- Continual Learning with Scaled Gradient ProjectionGobinda Saha, Kaushik RoyAAAI 2023 · 44 citations
Builds on9
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 295 citations
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 211 citations
Related papers
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
- Preserving Linear Separability in Continual Learning by Backward Feature ProjectionQiao Gu, Dongsub Shim, Florian ShkurtiCVPR 2023
- Continual Learning with Node-Importance based Adaptive Group Sparse RegularizationSangwon Jung, Hongjoon Ahn, Sungmin Cha, Taesup MoonNeurIPS 2020 · 176 citations
- CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningMarco Paul E. Apolinario, Sakshi Choudhary, Kaushik RoyICCV 2025 · 7 citations
- Natural continual learning: success is a journey, not (just) a destinationTa-Chu Kao, Kristopher T. Jensen, Gido van de Ven, Alberto Bernacchia et al.NeurIPS 2021 · 72 citations
