A Statistical Theory of Regularization-Based Continual Learning
Xuyang Zhao, Huiyuan Wang, Weiran Huang, Wei Lin
摘要
We provide a statistical analysis of regularization-based continual learning on a sequence of linear regression tasks, with emphasis on how different regularization terms affect the model performance. We first derive the convergence rate for the oracle estimator obtained as if all data were available simultaneously. Next, we consider a family of generalized -regularization algorithms indexed by matrix-valued hyperparameters, which includes the minimum norm estimator and continual ridge regression as special cases. As more tasks are introduced, we derive an iterative update formula for the estimation error of generalized -regularized estimators, from which we determine the hyperparameters resulting in the optimal algorithm. Interestingly, the choice of hyperparameters can effectively balance the trade-off between forward and backward knowledge transfer and adjust for data heterogeneity. Moreover, the estimation error of the optimal algorithm is derived explicitly, which is of the same order as that of the oracle estimator. In contrast, our lower bounds for the minimum norm estimator and continual ridge regression show their suboptimality. A byproduct of our theoretical analysis is the equivalence between early stopping and generalized -regularization in continual learning, which may be of independent interest. Finally, we conduct experiments to complement our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 被引用 25 次
- Provable Contrastive Continual LearningYichen Wen, Zhiquan Tan, Kaipeng Zheng, Chuanlong Xie 等ICML 2024 · 被引用 13 次
- Optimal Rates in Continual Linear Regression via Increasing RegularizationRan Levinstein, Amit Attia, Matan Schliserman, Uri Sherman 等NeurIPS 2025 · 被引用 10 次
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong 等NeurIPS 2025 · 被引用 5 次
- Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelsKai Tong, Kang Pan, Xiao Zhang, Erli Meng 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper12
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 被引用 206 次
- Gradient-based Editing of Memory Examples for Online Task-free Continual LearningXisen Jin, Arka Sadhu, Junyi Du, Xiang RenNeurIPS 2021 · 被引用 124 次
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 被引用 98 次
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 被引用 74 次
相关 Paper
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 被引用 4 次
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 被引用 23 次
- Last Iterate Convergence of Incremental Methods as a Model of ForgettingXufeng Cai, Jelena DiakonikolasICLR 2025
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 被引用 16 次
- Convergence and Implicit Bias of Gradient Descent on Continual Linear ClassificationHyunji Jung, Hanseul Cho, Chulhee YunICLR 2025
