Beyond Not-Forgetting: Continual Learning with Backward Knowledge Transfer
Sen Lin, Li Yang, Deliang Fan, Junshan Zhang
摘要
By learning a sequence of tasks continually, an agent in continual learning (CL) can improve the learning performance of both a new task and `old' tasks by leveraging the forward knowledge transfer and the backward knowledge transfer, respectively. However, most existing CL methods focus on addressing catastrophic forgetting in neural networks by minimizing the modification of the learnt model for old tasks. This inevitably limits the backward knowledge transfer from the new task to the old tasks, because judicious model updates could possibly improve the learning performance of the old tasks as well. To tackle this problem, we first theoretically analyze the conditions under which updating the learnt model of old tasks could be beneficial for CL and also lead to backward knowledge transfer, based on the gradient projection onto the input subspaces of old tasks. Building on the theoretical analysis, we next develop a ContinUal learning method with Backward knowlEdge tRansfer (CUBER), for a fixed capacity neural network without data replay. In particular, CUBER first characterizes the task correlation to identify the positively correlated old tasks in a layer-wise manner, and then selectively modifies the learnt model of the old tasks when learning the new task. Experimental studies show that CUBER can even achieve positive backward knowledge transfer on several existing CL benchmarks for the first time without data replay, where the related baselines still suffer from catastrophic forgetting (negative backward knowledge transfer). The superior performance of CUBER on the backward knowledge transfer also leads to higher accuracy accordingly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 被引用 74 次
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 被引用 39 次
- Learnability and Algorithm for Continual LearningGyuhak Kim, Changnan Xiao, Tatsuya Konishi, Bing LiuICML 2023 · 被引用 37 次
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 被引用 36 次
- Disentangling and mitigating the impact of task similarity for continual learningNaoki HirataniNeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper5
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 被引用 206 次
- Continual Learning of a Mixed Sequence of Similar and Dissimilar TasksZixuan Ke, Bing Liu, Xingchang HuangNeurIPS 2020 · 被引用 173 次
- TRGP: Trust Region Gradient Projection for Continual LearningSen Lin, Li Yang, Deliang Fan, Junshan ZhangICLR 2022 · 被引用 107 次
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 被引用 52 次
相关 Paper
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu 等CVPR 2021
- CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningMarco Paul E. Apolinario, Sakshi Choudhary, Kaushik RoyICCV 2025 · 被引用 7 次
- Introducing Common Null Space of Gradients for Gradient Projection Methods in Continual LearningChengyi Yang, Mingda Dong, Xiaoyue Zhang, Jiayin Qi 等ACM MM 2024 · 被引用 1 次
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng 等ICML 2024 · 被引用 23 次
- Turning Back Without Forgetting: Selective Backward Refinement for Parameter-Efficient Continual LearningAnushka Tiwari, Kaiyi JiICML 2026
