Layerwise Optimization by Gradient Decomposition for Continual Learning
Shixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu, Wanli Ouyang
摘要
Deep neural networks achieve state-of-the-art and sometimes super-human performance across various domains. However, when learning tasks sequentially, the networks easily forget the knowledge of previous tasks, known as "catastrophic forgetting". To achieve the consistencies between the old tasks and the new task, one effective solution is to modify the gradient for update. Previous methods enforce independent gradient constraints for different tasks, while we consider these gradients contain complex information, and propose to leverage inter-task information by gradient decomposition. In particular, the gradient of an old task is decomposed into a part shared by all old tasks and a part specific to that task. The gradient for update should be close to the gradient of the new task, consistent with the gradients shared by all old tasks, and orthogonal to the space spanned by the gradients specific to the old tasks. In this way, our approach encourages common knowledge consolidation without impairing the task-specific knowledge. Furthermore, the optimization is performed for the gradients of each layer separately rather than the concatenation of all gradients as in previous works. This effectively avoids the influence of the magnitude variation of the gradients in different layers. Extensive experiments validate the effectiveness of both gradient-decomposed optimization and layerwise updates. Our proposed method achieves state-of-theart results on various benchmarks of continual learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 被引用 397 次
- Continual Learning with Lifelong Vision TransformerZhen Wang, Liu Liu, Yiqun Duan, Yajing Kong 等CVPR 2022 · 被引用 63 次
- Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental LearningYujun Shi, Kuangqi Zhou, Jian Liang, Zihang Jiang 等CVPR 2022 · 被引用 57 次
- ICICLE: Interpretable Class Incremental Continual LearningDawid Rymarczyk, Joost van de Weijer, Bartosz Zielinski, Bartlomiej TwardowskiICCV 2023 · 被引用 35 次
- Interactive Continual Learning: Fast and Slow ThinkingBiqing Qi, Xinquan Chen, Junqi Gao, Dong Li 等CVPR 2024 · 被引用 15 次
它引用的顶会 Paper9
- Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-IDYixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao 等NeurIPS 2020 · 被引用 688 次
- Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identificationYixiao Ge, Dapeng Chen, Hongsheng LiICLR 2020 · 被引用 651 次
- Overcoming Catastrophic Forgetting With Unlabeled Data in the WildKibok Lee, Kimin Lee, Jinwoo Shin, Honglak LeeICCV 2019 · 被引用 231 次
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin 等ICLR 2020 · 被引用 221 次
- Gradient Regularized Contrastive Learning for Continual Domain AdaptationShixiang Tang, Peng Su, Dapeng Chen, Wanli OuyangAAAI 2021 · 被引用 72 次
相关 Paper
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 被引用 7 次
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 被引用 52 次
- Continual Learning with Scaled Gradient ProjectionGobinda Saha, Kaushik RoyAAAI 2023 · 被引用 44 次
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 被引用 98 次
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu 等ICCV 2023 · 被引用 28 次
