A Statistical Theory of Regularization-Based Continual Learning
Xuyang Zhao, Huiyuan Wang, Weiran Huang, Wei Lin
Abstract
We provide a statistical analysis of regularization-based continual learning on a sequence of linear regression tasks, with emphasis on how different regularization terms affect the model performance. We first derive the convergence rate for the oracle estimator obtained as if all data were available simultaneously. Next, we consider a family of generalized -regularization algorithms indexed by matrix-valued hyperparameters, which includes the minimum norm estimator and continual ridge regression as special cases. As more tasks are introduced, we derive an iterative update formula for the estimation error of generalized -regularized estimators, from which we determine the hyperparameters resulting in the optimal algorithm. Interestingly, the choice of hyperparameters can effectively balance the trade-off between forward and backward knowledge transfer and adjust for data heterogeneity. Moreover, the estimation error of the optimal algorithm is derived explicitly, which is of the same order as that of the oracle estimator. In contrast, our lower bounds for the minimum norm estimator and continual ridge regression show their suboptimality. A byproduct of our theoretical analysis is the equivalence between early stopping and generalized -regularization in continual learning, which may be of independent interest. Finally, we conduct experiments to complement our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06064b6e-da5b-4ab9-86f0-99634ac31fb2Cited by top-tier papers16
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 25 citations
- Provable Contrastive Continual LearningYichen Wen, Zhiquan Tan, Kaipeng Zheng, Chuanlong Xie et al.ICML 2024 · 13 citations
- Optimal Rates in Continual Linear Regression via Increasing RegularizationRan Levinstein, Amit Attia, Matan Schliserman, Uri Sherman et al.NeurIPS 2025 · 10 citations
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong et al.NeurIPS 2025 · 5 citations
- Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelsKai Tong, Kang Pan, Xiao Zhang, Erli Meng et al.ICCV 2025 · 3 citations
Builds on12
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 206 citations
- Gradient-based Editing of Memory Examples for Online Task-free Continual LearningXisen Jin, Arka Sadhu, Junyi Du, Xiang RenNeurIPS 2021 · 124 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 74 citations
Related papers
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 4 citations
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
- Last Iterate Convergence of Incremental Methods as a Model of ForgettingXufeng Cai, Jelena DiakonikolasICLR 2025
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
- Convergence and Implicit Bias of Gradient Descent on Continual Linear ClassificationHyunji Jung, Hanseul Cho, Chulhee YunICLR 2025
