Lune

ICLR2022Top-tier venue

Effect of scale on catastrophic forgetting in neural networks

Vinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan Dyer

2022Year
212Citations
58Top-tier citations

Abstract

Catastrophic forgetting presents a challenge in developing deep learning models capable of continual learning, i.e. learning tasks sequentially. Recently, both computer vision and natural-language processing have witnessed great progress through the use of large-scale pretrained models. In this work, we present an empirical study of catastrophic forgetting in this pretraining paradigm.Our experiments indicate that large, pretrained ResNets and Transformers are significantly more resistant to forgetting than randomly-initialized, trained-from-scratch models; this robustness systematically improves with scale of both model and pretraining dataset size.We take initial steps towards characterizing what aspect of model representations allows them to perform continual learning so well, finding that in the pretrained models, distinct class representations grow more orthogonal with scale. Our results suggest that, when possible, scale and a diverse pretraining dataset can be useful ingredients in mitigating catastrophic forgetting.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 8b41a02f-fd37-4555-8d4d-66c4aceefc0a

Cited by top-tier papers58

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines