Effective Continual Learning for Text Classification with Lightweight Snapshots
Jue Wang, Dajie Dong, Lidan Shou, Ke Chen, Gang Chen
摘要
Continual learning is known for suffering from catastrophic forgetting, a phenomenon where previously learned concepts are forgotten upon learning new tasks. A natural remedy is to use trained models for old tasks as ‘teachers’ to regularize the update of the current model to prevent such forgetting. However, this requires storing all past models, which is very space-consuming for large models, e.g. BERT, thus impractical in real-world applications. To tackle this issue, we propose to construct snapshots of seen tasks whose key knowledge is captured in lightweight adapters. During continual learning, we transfer knowledge from past snapshots to the current model through knowledge distillation, allowing the current model to review previously learned knowledge while learning new tasks. We also design representation recalibration to better handle the class-incremental setting. Experiments over various task sequences show that our approach effectively mitigates catastrophic forgetting and outperforms all baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
- Always Be Dreaming: A New Approach for Data-Free Class-Incremental LearningJames Seale Smith, Yen-Chang Hsu, Jonathan C. Balloch, Yilin Shen 等ICCV 2021 · 被引用 208 次
- Efficient Meta Lifelong-Learning with Limited MemoryZirui Wang, Sanket Vaibhav Mehta, Barnabás Póczos, Jaime G. CarbonellEMNLP 2020 · 被引用 47 次
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 被引用 36 次
- CPR: Classifier-Projection Regularization for Continual LearningSungmin Cha, Hsiang Hsu, Taebaek Hwang, Flávio P. Calmon 等ICLR 2021 · 被引用 32 次
相关 Paper
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 被引用 212 次
- Prototype-Sample Relation Distillation: Towards Replay-Free Continual LearningNader Asadi, MohammadReza Davari, Sudhir P. Mudur, Rahaf Aljundi 等ICML 2023 · 被引用 61 次
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal 等NeurIPS 2022 · 被引用 75 次
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding 等AAAI 2026
- Preserving Linear Separability in Continual Learning by Backward Feature ProjectionQiao Gu, Dongsub Shim, Florian ShkurtiCVPR 2023
