Understanding the Dynamics of Forgetting and Generalization in Continual Learning via the Neural Tangent Kernel
Guodong Zheng, Peng Wang, Shengchao Hu, Quan Zheng, Li Shen
Abstract
Continual learning (CL) enables models to acquire new tasks sequentially while retaining previously learned knowledge. However, most theoretical analyses focus on simplified, converged models or restrictive data distributions and therefore fail to capture how forgetting and generalization evolve during training in more general settings. Current theory faces two fundamental challenges: (i) analyses confined to the converged regime cannot characterize intermediate training dynamics; and (ii) establishing forgetting bounds requires two-sided bounds on the population risk for each task. To address these challenges, we analyze the training-time dynamics of forgetting and generalization in standard CL within the Neural Tangent Kernel (NTK) regime, showing that decreasing the loss’s Lipschitz constant and minimizing the cross-task kernel jointly reduce forgetting and improve generalization. Specifically, we (i) characterize intermediate training stages via kernel gradient flow and (ii) employ Rademacher complexity to derive both upper and lower bounds on population risk. Building on these insights, we propose OGD+, which projects the current task’s gradient onto the orthogonal complement of the subspace spanned by gradients of the most recent task evaluated on all prior samples. We further introduce Orthogonal Penalized Gradient Descent (OPGD), which augments OGD+ with gradient-norm penalization to jointly reduce forgetting and enhance generalization. Experiments on multiple benchmarks corroborate our theoretical predictions and demonstrate the effectiveness of OPGD, providing a principled pathway from theory to algorithm design in CL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a52091d6-0396-4fc7-b040-5a7010f627abBuilds on17
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Penalizing Gradient Norm for Efficiently Improving Generalization in Deep LearningYang Zhao, Hao Zhang, Xiuyuan HuICML 2022 · 165 citations
- TRGP: Trust Region Gradient Projection for Continual LearningSen Lin, Li Yang, Deliang Fan, Junshan ZhangICLR 2022 · 107 citations
Related papers
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
- Optimal Rates in Continual Linear Regression via Increasing RegularizationRan Levinstein, Amit Attia, Matan Schliserman, Uri Sherman et al.NeurIPS 2025 · 10 citations
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 4 citations
- Convergence and Implicit Bias of Gradient Descent on Continual Linear ClassificationHyunji Jung, Hanseul Cho, Chulhee YunICLR 2025
