Continual Learning with Scaled Gradient Projection
Gobinda Saha, Kaushik Roy
Abstract
In neural networks, continual learning results in gradient interference among sequential tasks, leading to catastrophic forgetting of old tasks while learning new ones. This issue is addressed in recent methods by storing the important gradient spaces for old tasks and updating the model orthogonally during new tasks. However, such restrictive orthogonal gradient updates hamper the learning capability of the new tasks resulting in sub-optimal performance. To improve new learning while minimizing forgetting, in this paper we propose a Scaled Gradient Projection (SGP) method, where we combine the orthogonal gradient projections with scaled gradient steps along the important gradient spaces for the past tasks. The degree of gradient scaling along these spaces depends on the importance of the bases spanning them. We propose an efficient method for computing and accumulating importance of these bases using the singular value decomposition of the input representations for each task. We conduct extensive experiments ranging from continual image classification to reinforcement learning tasks and report better performance with less training overhead than the state-of-the-art approaches. Codes: https://github.com/sahagobinda/SGP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Evolving Standardization for Continual Domain Generalization over Temporal DriftMixue Xie, Shuang Li, Longhui Yuan, Chi Harold Liu et al.NeurIPS 2023 · 21 citations
- Layerwise Proximal Replay: A Proximal Point Method for Online Continual LearningJinsoo Yoo, Yunpeng Liu, Frank Wood, Geoff PleissICML 2024 · 13 citations
- Continual Gradient Low-Rank Projection Fine-Tuning for LLMsChenxu Wang, Yilin Lyu, Zicheng Sun, Liping JingACL 2025 · 7 citations
- CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningMarco Paul E. Apolinario, Sakshi Choudhary, Kaushik RoyICCV 2025 · 7 citations
- TinySubNets: An Efficient and Low Capacity Continual Learning StrategyMarcin Pietron, Kamil Faber, Dominik Zurek, Roberto CorizzoAAAI 2025 · 6 citations
Builds on14
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Using Hindsight to Anchor Past Knowledge in Continual LearningArslan Chaudhry, Albert Gordo, Puneet K. Dokania, Philip H. S. Torr et al.AAAI 2021 · 279 citations
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 211 citations
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 206 citations
Related papers
- TRGP: Trust Region Gradient Projection for Continual LearningSen Lin, Li Yang, Deliang Fan, Junshan ZhangICLR 2022 · 107 citations
- Class Gradient Projection For Continual LearningCheng Chen, Ji Zhang, Jingkuan Song, Lianli GaoACM MM 2022 · 14 citations
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
- Adaptive Orthogonal Projection for Batch and Online Continual LearningYiduo Guo, Wenpeng Hu, Dongyan Zhao, Bing LiuAAAI 2022 · 56 citations
- Introducing Common Null Space of Gradients for Gradient Projection Methods in Continual LearningChengyi Yang, Mingda Dong, Xiaoyue Zhang, Jiayin Qi et al.ACM MM 2024 · 1 citation
