Beyond Not-Forgetting: Continual Learning with Backward Knowledge Transfer
Sen Lin, Li Yang, Deliang Fan, Junshan Zhang
Abstract
By learning a sequence of tasks continually, an agent in continual learning (CL) can improve the learning performance of both a new task and `old' tasks by leveraging the forward knowledge transfer and the backward knowledge transfer, respectively. However, most existing CL methods focus on addressing catastrophic forgetting in neural networks by minimizing the modification of the learnt model for old tasks. This inevitably limits the backward knowledge transfer from the new task to the old tasks, because judicious model updates could possibly improve the learning performance of the old tasks as well. To tackle this problem, we first theoretically analyze the conditions under which updating the learnt model of old tasks could be beneficial for CL and also lead to backward knowledge transfer, based on the gradient projection onto the input subspaces of old tasks. Building on the theoretical analysis, we next develop a ContinUal learning method with Backward knowlEdge tRansfer (CUBER), for a fixed capacity neural network without data replay. In particular, CUBER first characterizes the task correlation to identify the positively correlated old tasks in a layer-wise manner, and then selectively modifies the learnt model of the old tasks when learning the new task. Experimental studies show that CUBER can even achieve positive backward knowledge transfer on several existing CL benchmarks for the first time without data replay, where the related baselines still suffer from catastrophic forgetting (negative backward knowledge transfer). The superior performance of CUBER on the backward knowledge transfer also leads to higher accuracy accordingly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42d47488-b02c-44e4-99e2-7ee7a9cfec70Cited by top-tier papers36
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 74 citations
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 39 citations
- Learnability and Algorithm for Continual LearningGyuhak Kim, Changnan Xiao, Tatsuya Konishi, Bing LiuICML 2023 · 37 citations
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Disentangling and mitigating the impact of task similarity for continual learningNaoki HirataniNeurIPS 2024 · 21 citations
Builds on5
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 206 citations
- Continual Learning of a Mixed Sequence of Similar and Dissimilar TasksZixuan Ke, Bing Liu, Xingchang HuangNeurIPS 2020 · 173 citations
- TRGP: Trust Region Gradient Projection for Continual LearningSen Lin, Li Yang, Deliang Fan, Junshan ZhangICLR 2022 · 107 citations
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 52 citations
Related papers
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu et al.CVPR 2021
- CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningMarco Paul E. Apolinario, Sakshi Choudhary, Kaushik RoyICCV 2025 · 7 citations
- Introducing Common Null Space of Gradients for Gradient Projection Methods in Continual LearningChengyi Yang, Mingda Dong, Xiaoyue Zhang, Jiayin Qi et al.ACM MM 2024 · 1 citation
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng et al.ICML 2024 · 23 citations
- Turning Back Without Forgetting: Selective Backward Refinement for Parameter-Efficient Continual LearningAnushka Tiwari, Kaiyi JiICML 2026
