Layerwise Optimization by Gradient Decomposition for Continual Learning
Shixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu, Wanli Ouyang
Abstract
Deep neural networks achieve state-of-the-art and sometimes super-human performance across various domains. However, when learning tasks sequentially, the networks easily forget the knowledge of previous tasks, known as "catastrophic forgetting". To achieve the consistencies between the old tasks and the new task, one effective solution is to modify the gradient for update. Previous methods enforce independent gradient constraints for different tasks, while we consider these gradients contain complex information, and propose to leverage inter-task information by gradient decomposition. In particular, the gradient of an old task is decomposed into a part shared by all old tasks and a part specific to that task. The gradient for update should be close to the gradient of the new task, consistent with the gradients shared by all old tasks, and orthogonal to the space spanned by the gradients specific to the old tasks. In this way, our approach encourages common knowledge consolidation without impairing the task-specific knowledge. Furthermore, the optimization is performed for the gradients of each layer separately rather than the concatenation of all gradients as in previous works. This effectively avoids the influence of the magnitude variation of the gradients in different layers. Extensive experiments validate the effectiveness of both gradient-decomposed optimization and layerwise updates. Our proposed method achieves state-of-theart results on various benchmarks of continual learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba424bbc-fe23-443a-b470-61fdf93d170fCited by top-tier papers15
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 397 citations
- Continual Learning with Lifelong Vision TransformerZhen Wang, Liu Liu, Yiqun Duan, Yajing Kong et al.CVPR 2022 · 63 citations
- Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental LearningYujun Shi, Kuangqi Zhou, Jian Liang, Zihang Jiang et al.CVPR 2022 · 57 citations
- ICICLE: Interpretable Class Incremental Continual LearningDawid Rymarczyk, Joost van de Weijer, Bartosz Zielinski, Bartlomiej TwardowskiICCV 2023 · 35 citations
- Interactive Continual Learning: Fast and Slow ThinkingBiqing Qi, Xinquan Chen, Junqi Gao, Dong Li et al.CVPR 2024 · 15 citations
Builds on9
- Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-IDYixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao et al.NeurIPS 2020 · 688 citations
- Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identificationYixiao Ge, Dapeng Chen, Hongsheng LiICLR 2020 · 651 citations
- Overcoming Catastrophic Forgetting With Unlabeled Data in the WildKibok Lee, Kimin Lee, Jinwoo Shin, Honglak LeeICCV 2019 · 231 citations
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
- Gradient Regularized Contrastive Learning for Continual Domain AdaptationShixiang Tang, Peng Su, Dapeng Chen, Wanli OuyangAAAI 2021 · 72 citations
Related papers
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 7 citations
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 52 citations
- Continual Learning with Scaled Gradient ProjectionGobinda Saha, Kaushik RoyAAAI 2023 · 44 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
