Does Continual Learning Equally Forget All Parameters?
Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang, Chengqi Zhang
摘要
Distribution shift (e.g., task or domain shift) in continual learning (CL) usually results in catastrophic forgetting of neural networks. Although it can be alleviated by repeatedly replaying buffered data, the every-step replay is time-consuming. In this paper, we study which modules in neural networks are more prone to forgetting by investigating their training dynamics during CL. Our proposed metrics show that only a few modules are more task-specific and sensitively alter between tasks, while others can be shared across tasks as common knowledge. Hence, we attribute forgetting mainly to the former and find that finetuning them only on a small buffer at the end of any CL method can bring non-trivial improvement. Due to the small number of finetuned parameters, such Forgetting Prioritized Finetuning (FPF)'' is efficient in computation. We further propose a more efficient and simpler method that entirely removes the every-step replay and replaces them by only $k$-times of FPF periodically triggered during CL. Surprisingly, this -FPF'' performs comparably to FPF and outperforms the SOTA CL methods but significantly reduces their computational overhead and cost. In experiments on several benchmarks of class- and domain-incremental CL, FPF consistently improves existing CL methods by a large margin, and -FPF further excels in efficiency without degrading the accuracy. We also empirically studied the impact of buffer size, epochs per task, and finetuning modules on the cost and accuracy of our methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Continual Task Allocation in Meta-Policy Network via Sparse PromptingYijun Yang, Tianyi Zhou, Jing Jiang, Guodong Long 等ICML 2023 · 被引用 14 次
- Mitigate Catastrophic Remembering via Continual Knowledge Purification for Noisy Lifelong Person Re-IdentificationKunlun Xu, Haozhuo Zhang, Yu Li, Yuxin Peng 等ACM MM 2024 · 被引用 10 次
- Demystifying Language Model Forgetting with Low-rank Example AssociationsXisen Jin, Xiang RenNeurIPS 2025 · 被引用 9 次
- A Layer Selection Approach to Test Time AdaptationSabyasachi Sahoo, Mostafa ElAraby, Jonas Ngnawé, Yann Batiste Pequignot 等AAAI 2025 · 被引用 6 次
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao 等EMNLP 2024 · 被引用 4 次
它引用的顶会 Paper4
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 被引用 207 次
- Pretrained Language Model in Continual Learning: A Comparative StudyTongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li 等ICLR 2022 · 被引用 76 次
- Continual Normalization: Rethinking Batch Normalization for Online Continual LearningQuang Pham, Chenghao Liu, Steven C. H. HoiICLR 2022 · 被引用 72 次
- Continual Adaptation of Visual Representations via Domain Randomization and Meta-LearningRiccardo Volpi, Diane Larlus, Grégory RogezCVPR 2021
相关 Paper
- Predicting the Susceptibility of Examples to Catastrophic ForgettingGuy Hacohen, Tinne TuytelaarsICML 2025
- Probing Representation Forgetting in Supervised and Unsupervised Continual LearningMohammadReza Davari, Nader Asadi, Sudhir P. Mudur, Rahaf Aljundi 等CVPR 2022 · 被引用 48 次
- On Generalizing Beyond Domains in Cross-Domain Continual LearningChristian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu 等CVPR 2022 · 被引用 34 次
- Retrospective Adversarial Replay for Continual LearningLilly Kumari, Shengjie Wang, Tianyi Zhou, Jeff A. BilmesNeurIPS 2022 · 被引用 57 次
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng 等ICML 2024 · 被引用 23 次
