Effect of scale on catastrophic forgetting in neural networks
Vinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan Dyer
摘要
Catastrophic forgetting presents a challenge in developing deep learning models capable of continual learning, i.e. learning tasks sequentially. Recently, both computer vision and natural-language processing have witnessed great progress through the use of large-scale pretrained models. In this work, we present an empirical study of catastrophic forgetting in this pretraining paradigm.Our experiments indicate that large, pretrained ResNets and Transformers are significantly more resistant to forgetting than randomly-initialized, trained-from-scratch models; this robustness systematically improves with scale of both model and pretraining dataset size.We take initial steps towards characterizing what aspect of model representations allows them to perform continual learning so well, finding that in the pretrained models, distinct class representations grow more orthogonal with scale. Our results suggest that, when possible, scale and a diverse pretraining dataset can be useful ingredients in mitigating catastrophic forgetting.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper58
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 被引用 304 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song 等NeurIPS 2022 · 被引用 230 次
- SLCA: Slow Learner with Classifier Alignment for Continual Learning on a Pre-trained ModelGengwei Zhang, Liyuan Wang, Guoliang Kang, Ling Chen 等ICCV 2023 · 被引用 196 次
- Hierarchical Decomposition of Prompt-Based Continual Learning: Rethinking Obscured Sub-optimalityLiyuan Wang, Jingyi Xie, Xingxing Zhang, Mingyi Huang 等NeurIPS 2023 · 被引用 183 次
相关 Paper
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual LearningHuihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu 等ICML 2026 · 被引用 14 次
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen 等AAAI 2023 · 被引用 4 次
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal 等NeurIPS 2022 · 被引用 75 次
- Pretrained Language Model in Continual Learning: A Comparative StudyTongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li 等ICLR 2022 · 被引用 76 次
- On the Diminishing Returns of Width for Continual LearningEtash Kumar Guha, Vihan LakshmanICML 2024 · 被引用 9 次
