Learning curves for continual learning in neural networks: Self-knowledge transfer and forgetting
Ryo Karakida, Shotaro Akaho
摘要
Sequential training from task to task is becoming one of the major objects in deep learning applications such as continual learning and transfer learning. Nevertheless, it remains unclear under what conditions the trained model's performance improves or deteriorates. To deepen our understanding of sequential training, this study provides a theoretical analysis of generalization performance in a solvable case of continual learning. We consider neural networks in the neural tangent kernel (NTK) regime that continually learn target functions from task to task, and investigate the generalization by using an established statistical mechanical analysis of kernel ridge-less regression. We first show characteristic transitions from positive to negative transfer. More similar targets above a specific critical value can achieve positive knowledge transfer for the subsequent task while catastrophic forgetting occurs even with very similar targets. Next, we investigate a variant of continual learning which supposes the same target function in multiple tasks. Even for the same target, the trained model shows some transfer and forgetting depending on the sample size of each task. We can guarantee that the generalization error monotonically decreases from task to task for equal sample sizes while unbalanced sample sizes deteriorate the generalization. We respectively refer to these improvement and deterioration as self-knowledge transfer and forgetting, and empirically confirm them in realistic training of deep neural networks as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Theoretical Study on Solving Continual LearningGyuhak Kim, Changnan Xiao, Tatsuya Konishi, Zixuan Ke 等NeurIPS 2022 · 被引用 119 次
- Disentangling and mitigating the impact of task similarity for continual learningNaoki HirataniNeurIPS 2024 · 被引用 21 次
- What Will My Model Forget? Forecasting Forgotten Examples in Language Model RefinementXisen Jin, Xiang RenICML 2024 · 被引用 8 次
- Understanding the Dynamics of Forgetting and Generalization in Continual Learning via the Neural Tangent KernelGuodong Zheng, Peng Wang, Shengchao Hu, Quan Zheng 等ICLR 2026
- Graceful Forgetting in Generative Language ModelsChunyang Jiang, Chi-Min Chan, Yiyang Cai, Yulong Liu 等EMNLP 2025
它引用的顶会 Paper11
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 被引用 207 次
相关 Paper
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 被引用 98 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 被引用 23 次
- BNS: Building Network Structures Dynamically for Continual LearningQi Qin, Wenpeng Hu, Han Peng, Dongyan Zhao 等NeurIPS 2021 · 被引用 54 次
- On the Diminishing Returns of Width for Continual LearningEtash Kumar Guha, Vihan LakshmanICML 2024 · 被引用 9 次
