On the Diminishing Returns of Width for Continual Learning
Etash Kumar Guha, Vihan Lakshman
摘要
While deep neural networks have demonstrated groundbreaking performance in various settings, these models often suffer from catastrophic forgetting when trained on new tasks in sequence. Several works have empirically demonstrated that increasing the width of a neural network leads to a decrease in catastrophic forgetting but have yet to characterize the exact relationship between width and continual learning. We design one of the first frameworks to analyze Continual Learning Theory and prove that width is directly related to forgetting in Feed-Forward Networks (FFN). Specifically, we demonstrate that increasing network widths to reduce forgetting yields diminishing returns. We empirically verify our claims at widths hitherto unexplored in prior studies where the diminishing returns are clearly observed as predicted by our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Measuring Representational Shifts in Continual Learning: A Linear Transformation PerspectiveJoonkyu Kim, Yejin Kim, Jy-yong SohnICML 2025
- LoRanPAC: Low-rank Random Features and Pre-trained Models for Bridging Theory and Practice in Continual LearningLiangzu Peng, Juan Elenter, Joshua Agterberg, Alejandro Ribeiro 等ICLR 2025
它引用的顶会 Paper9
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 被引用 323 次
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 被引用 212 次
- Linear Mode Connectivity in Multitask and Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Dilan Görür, Razvan Pascanu 等ICLR 2021 · 被引用 176 次
- A Theoretical Study on Solving Continual LearningGyuhak Kim, Changnan Xiao, Tatsuya Konishi, Zixuan Ke 等NeurIPS 2022 · 被引用 119 次
- Optimal Continual Learning has Perfect Memory and is NP-hardJeremias Knoblauch, Hisham Husain, Tom DietheICML 2020 · 被引用 116 次
相关 Paper
- Wide Neural Networks Forget Less CatastrophicallySeyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu 等ICML 2022 · 被引用 84 次
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann 等ICML 2025
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 被引用 74 次
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 被引用 98 次
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 被引用 23 次
