Wide Neural Networks Forget Less Catastrophically
Seyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu, Razvan Pascanu, Dilan Görür, Mehrdad Farajtabar
Abstract
A primary focus area in continual learning research is alleviating the "catastrophic forgetting" problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual learning literature is encouraging, our understanding of what properties of neural networks contribute to catastrophic forgetting is still limited. To address this, instead of focusing on continual learning algorithms, in this work, we focus on the model itself and study the impact of "width" of the neural network architecture on catastrophic forgetting, and show that width has a surprisingly significant effect on forgetting. To explain this effect, we study the learning dynamics of the network from various perspectives such as gradient orthogonality, sparsity, and lazy training regime. We provide potential explanations that are consistent with the empirical results across different architectures and continual learning benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 304 citations
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu et al.NeurIPS 2023 · 40 citations
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 39 citations
- Learning and Forgetting Unsafe Examples in Large Language ModelsJiachen Zhao, Zhun Deng, David Madras, James Zou et al.ICML 2024 · 27 citations
- The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting - An Analytical ModelDaniel Goldfarb, Itay Evron, Nir Weinberger, Daniel Soudry et al.ICLR 2024 · 25 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 295 citations
- Using Hindsight to Anchor Past Knowledge in Continual LearningArslan Chaudhry, Albert Gordo, Puneet K. Dokania, Philip H. S. Torr et al.AAAI 2021 · 279 citations
Related papers
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann et al.ICML 2025
- On the Diminishing Returns of Width for Continual LearningEtash Kumar Guha, Vihan LakshmanICML 2024 · 9 citations
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang et al.ICML 2023 · 21 citations
- On the Theory of Continual Learning with Gradient Descent for Neural NetworksHossein Taheri, Avishek Ghosh, Arya MazumdarICML 2026 · 2 citations
