Memory Efficient Continual Learning with Transformers
Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, Cédric Archambeau
摘要
In many real-world scenarios, data to train machine learning models becomes available over time. Unfortunately, these models struggle to continually learn new concepts without forgetting what has been learnt in the past. This phenomenon is known as catastrophic forgetting and it is difficult to prevent due to practical constraints. For instance, the amount of data that can be stored or the computational resources that can be used might be limited. Moreover, applications increasingly rely on large pre-trained neural networks, such as pre-trained Transformers, since the resources or data might not be available in sufficiently large quantities to practitioners to train the model from scratch. In this paper, we devise a method to incrementally train a model on a sequence of tasks using pre-trained Transformers and extending them with Adapters. Different than the existing approaches, our method is able to scale to a large number of tasks without significant overhead and allows sharing information across tasks. On both image and text classification tasks, we empirically demonstrate that our method maintains a good predictive performance without retraining the model or increasing the number of model parameters over time. The resulting model is also significantly faster at inference time compared to Adapter-based state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- RanPAC: Random Projections and Pre-trained Models for Continual LearningMark D. McDonnell, Dong Gong, Amin Parvaneh, Ehsan Abbasnejad 等NeurIPS 2023 · 被引用 245 次
- Towards Modular LLMs by Building and Reusing a Library of LoRAsOleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin 等ICML 2024 · 被引用 70 次
- Semantically-Shifted Incremental Adapter-Tuning is A Continual ViTransformerYuwen Tan, Qinhao Zhou, Xiang Xiang, Ke Wang 等CVPR 2024 · 被引用 14 次
- On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of CodeMartin Weyssow, Xin Zhou, Kisub Kim, David Lo 等FSE 2023 · 被引用 9 次
- Prompts Can Play Lottery Tickets Well: Achieving Lifelong Information Extraction via Lottery Prompt TuningZujie Liang, Feng Wei, Yin Jie, Yuxi Qian 等ACL 2023 · 被引用 8 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- DyTox: Transformers for Continual Learning with DYnamic TOken eXpansionArthur Douillard, Alexandre Ramé, Guillaume Couairon, Matthieu CordCVPR 2022 · 被引用 315 次
相关 Paper
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 被引用 212 次
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen 等AAAI 2023 · 被引用 4 次
- Learn or Recall? Revisiting Incremental Learning with Pre-trained Language ModelsJunhao Zheng, Shengjie Qiu, Qianli MaACL 2024
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan 等AAAI 2025 · 被引用 31 次
- Can BERT Refrain from Forgetting on Sequential Tasks? A Probing StudyMingxu Tao, Yansong Feng, Dongyan ZhaoICLR 2023
