Memory-Statistics Tradeoff in Continual Learning with Structural Regularization
Haoran Li, Jingfeng Wu, Vladimir Braverman
Abstract
We study the statistical performance of a continual learning problem with two linear regression tasks in a well-specified random design setting. We consider a structural regularization algorithm that incorporates a generalized -regularization tailored to the Hessian of the previous task for mitigating catastrophic forgetting. We establish upper and lower bounds on the joint excess risk for this algorithm. Our analysis reveals a fundamental trade-off between memory complexity and statistical efficiency, where memory complexity is measured by the number of vectors needed to define the structural regularization. Specifically, increasing the number of vectors in structural regularization leads to a worse memory complexity but an improved excess risk, and vice versa. Furthermore, our theory suggests that naive continual learning without regularization suffers from catastrophic forgetting, while structural regularization mitigates this issue. Notably, structural regularization achieves comparable performance to joint training with access to both tasks simultaneously. These results highlight the critical role of curvature-aware regularization for continual learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e3b7ee0-9f9f-4b2b-bac9-2d4f76a0590dCited by top-tier papers4
- Optimal Rates in Continual Linear Regression via Increasing RegularizationRan Levinstein, Amit Attia, Matan Schliserman, Uri Sherman et al.NeurIPS 2025 · 10 citations
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong et al.NeurIPS 2025 · 5 citations
- Compact Memory for Continual Logistic RegressionYohan Jung, Hyungi Lee, Wenlong Chen, Thomas Möllenhoff et al.NeurIPS 2025 · 2 citations
- Understanding the Dynamics of Forgetting and Generalization in Continual Learning via the Neural Tangent KernelGuodong Zheng, Peng Wang, Shengchao Hu, Quan Zheng et al.ICLR 2026
Builds on6
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 74 citations
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
- Sliced Cramer Synaptic Consolidation for Preserving Deeply Learned RepresentationsSoheil Kolouri, Nicholas A. Ketz, Andrea Soltoggio, Praveen K. PillyICLR 2020 · 42 citations
- Continual Learning in Linear Classification on Separable DataItay Evron, Edward Moroshko, Gon Buzaglo, Maroun Khriesh et al.ICML 2023 · 32 citations
Related papers
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
- The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting - An Analytical ModelDaniel Goldfarb, Itay Evron, Nir Weinberger, Daniel Soudry et al.ICLR 2024 · 25 citations
- Optimizing Spca-based Continual Learning: A Theoretical ApproachChunchun Yang, Malik Tiomoko, Zengfu WangICLR 2023
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
- A Statistical Theory of Regularization-Based Continual LearningXuyang Zhao, Huiyuan Wang, Weiran Huang, Wei LinICML 2024 · 40 citations
