Wide Neural Networks Forget Less Catastrophically
Seyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu, Razvan Pascanu, Dilan Görür, Mehrdad Farajtabar
摘要
A primary focus area in continual learning research is alleviating the "catastrophic forgetting" problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual learning literature is encouraging, our understanding of what properties of neural networks contribute to catastrophic forgetting is still limited. To address this, instead of focusing on continual learning algorithms, in this work, we focus on the model itself and study the impact of "width" of the neural network architecture on catastrophic forgetting, and show that width has a surprisingly significant effect on forgetting. To explain this effect, we study the learning dynamics of the network from various perspectives such as gradient orthogonality, sparsity, and lazy training regime. We provide potential explanations that are consistent with the empirical results across different architectures and continual learning benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 被引用 304 次
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu 等NeurIPS 2023 · 被引用 40 次
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 被引用 39 次
- Learning and Forgetting Unsafe Examples in Large Language ModelsJiachen Zhao, Zhun Deng, David Madras, James Zou 等ICML 2024 · 被引用 27 次
- The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting - An Analytical ModelDaniel Goldfarb, Itay Evron, Nir Weinberger, Daniel Soudry 等ICLR 2024 · 被引用 25 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 被引用 295 次
- Using Hindsight to Anchor Past Knowledge in Continual LearningArslan Chaudhry, Albert Gordo, Puneet K. Dokania, Philip H. S. Torr 等AAAI 2021 · 被引用 279 次
相关 Paper
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann 等ICML 2025
- On the Diminishing Returns of Width for Continual LearningEtash Kumar Guha, Vihan LakshmanICML 2024 · 被引用 9 次
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 被引用 212 次
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang 等ICML 2023 · 被引用 21 次
- On the Theory of Continual Learning with Gradient Descent for Neural NetworksHossein Taheri, Avishek Ghosh, Arya MazumdarICML 2026 · 被引用 2 次
