On the Diminishing Returns of Width for Continual Learning
Etash Kumar Guha, Vihan Lakshman
Abstract
While deep neural networks have demonstrated groundbreaking performance in various settings, these models often suffer from catastrophic forgetting when trained on new tasks in sequence. Several works have empirically demonstrated that increasing the width of a neural network leads to a decrease in catastrophic forgetting but have yet to characterize the exact relationship between width and continual learning. We design one of the first frameworks to analyze Continual Learning Theory and prove that width is directly related to forgetting in Feed-Forward Networks (FFN). Specifically, we demonstrate that increasing network widths to reduce forgetting yields diminishing returns. We empirically verify our claims at widths hitherto unexplored in prior studies where the diminishing returns are clearly observed as predicted by our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd2b2520-387b-41a8-8df7-49d040640482Cited by top-tier papers2
- Measuring Representational Shifts in Continual Learning: A Linear Transformation PerspectiveJoonkyu Kim, Yejin Kim, Jy-yong SohnICML 2025
- LoRanPAC: Low-rank Random Features and Pre-trained Models for Bridging Theory and Practice in Continual LearningLiangzu Peng, Juan Elenter, Joshua Agterberg, Alejandro Ribeiro et al.ICLR 2025
Builds on9
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 323 citations
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
- Linear Mode Connectivity in Multitask and Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Dilan Görür, Razvan Pascanu et al.ICLR 2021 · 176 citations
- A Theoretical Study on Solving Continual LearningGyuhak Kim, Changnan Xiao, Tatsuya Konishi, Zixuan Ke et al.NeurIPS 2022 · 119 citations
- Optimal Continual Learning has Perfect Memory and is NP-hardJeremias Knoblauch, Hisham Husain, Tom DietheICML 2020 · 116 citations
Related papers
- Wide Neural Networks Forget Less CatastrophicallySeyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu et al.ICML 2022 · 84 citations
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann et al.ICML 2025
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 74 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
