Maslow's Hammer in Catastrophic Forgetting: Node Re-Use vs. Node Activation
Sebastian Lee, Stefano Sarao Mannelli, Claudia Clopath, Sebastian Goldt, Andrew M. Saxe
Abstract
Continual learning - learning new tasks in sequence while maintaining performance on old tasks - remains particularly challenging for artificial neural networks. Surprisingly, the amount of forgetting does not increase with the dissimilarity between the learned tasks, but appears to be worst in an intermediate similarity regime. In this paper we theoretically analyse both a synthetic teacher-student framework and a real data setup to provide an explanation of this phenomenon that we name Maslow's hammer hypothesis. Our analysis reveals the presence of a trade-off between node activation and node re-use that results in worst forgetting in the intermediate regime. Using this understanding we reinterpret popular algorithmic interventions for catastrophic interference in terms of this trade-off, and identify the regimes in which they are most effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fbf6aaf-184c-48c7-9987-7ff188746277Cited by top-tier papers3
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu et al.NeurIPS 2023 · 40 citations
- Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum LearningJin Hwa Lee, Stefano Sarao Mannelli, Andrew M. SaxeICML 2024 · 15 citations
- Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2025
Builds on8
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
- Efficient Continual Learning with Modular Networks and Task-Driven PriorsTom Veniat, Ludovic Denoyer, Marc'Aurelio RanzatoICLR 2021 · 110 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
- Continual Learning via Local Module CompositionOleksiy Ostapenko, Pau Rodríguez, Massimo Caccia, Laurent CharlinNeurIPS 2021 · 98 citations
- Wide Neural Networks Forget Less CatastrophicallySeyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu et al.ICML 2022 · 84 citations
Related papers
- A Theory of Initialisation's Impact on SpecialisationDevon Jarvis, Sebastian Lee, Clémentine Carla Juliette Dominé, Andrew M. Saxe et al.ICLR 2025
- The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting - An Analytical ModelDaniel Goldfarb, Itay Evron, Nir Weinberger, Daniel Soudry et al.ICLR 2024 · 25 citations
- Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual LearningSanghwan Kim, Lorenzo Noci, Antonio Orvieto, Thomas HofmannCVPR 2023
- Understanding Forgetting in Continual Learning with Linear RegressionMeng Ding, Kaiyi Ji, Di Wang, Jinhui XuICML 2024 · 23 citations
- Sequential Mastery of Multiple Visual Tasks: Networks Naturally Learn to Learn and Forget to ForgetGuy Davidson, Michael C. MozerCVPR 2020
