Memory Efficient Continual Learning with Transformers
Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, Cédric Archambeau
Abstract
In many real-world scenarios, data to train machine learning models becomes available over time. Unfortunately, these models struggle to continually learn new concepts without forgetting what has been learnt in the past. This phenomenon is known as catastrophic forgetting and it is difficult to prevent due to practical constraints. For instance, the amount of data that can be stored or the computational resources that can be used might be limited. Moreover, applications increasingly rely on large pre-trained neural networks, such as pre-trained Transformers, since the resources or data might not be available in sufficiently large quantities to practitioners to train the model from scratch. In this paper, we devise a method to incrementally train a model on a sequence of tasks using pre-trained Transformers and extending them with Adapters. Different than the existing approaches, our method is able to scale to a large number of tasks without significant overhead and allows sharing information across tasks. On both image and text classification tasks, we empirically demonstrate that our method maintains a good predictive performance without retraining the model or increasing the number of model parameters over time. The resulting model is also significantly faster at inference time compared to Adapter-based state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bd5b9ac-ba67-4971-9148-be42037ec782Cited by top-tier papers15
- RanPAC: Random Projections and Pre-trained Models for Continual LearningMark D. McDonnell, Dong Gong, Amin Parvaneh, Ehsan Abbasnejad et al.NeurIPS 2023 · 245 citations
- Towards Modular LLMs by Building and Reusing a Library of LoRAsOleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin et al.ICML 2024 · 70 citations
- Semantically-Shifted Incremental Adapter-Tuning is A Continual ViTransformerYuwen Tan, Qinhao Zhou, Xiang Xiang, Ke Wang et al.CVPR 2024 · 14 citations
- On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of CodeMartin Weyssow, Xin Zhou, Kisub Kim, David Lo et al.FSE 2023 · 9 citations
- Prompts Can Play Lottery Tickets Well: Achieving Lifelong Information Extraction via Lottery Prompt TuningZujie Liang, Feng Wei, Yin Jie, Yuxi Qian et al.ACL 2023 · 8 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- DyTox: Transformers for Continual Learning with DYnamic TOken eXpansionArthur Douillard, Alexandre Ramé, Guillaume Couairon, Matthieu CordCVPR 2022 · 315 citations
Related papers
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen et al.AAAI 2023 · 4 citations
- Learn or Recall? Revisiting Incremental Learning with Pre-trained Language ModelsJunhao Zheng, Shengjie Qiu, Qianli MaACL 2024
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan et al.AAAI 2025 · 31 citations
- Can BERT Refrain from Forgetting on Sequential Tasks? A Probing StudyMingxu Tao, Yansong Feng, Dongyan ZhaoICLR 2023
