Can BERT Refrain from Forgetting on Sequential Tasks? A Probing Study
Mingxu Tao, Yansong Feng, Dongyan Zhao
摘要
Large pre-trained language models help to achieve state of the art on a variety of natural language processing (NLP) tasks, nevertheless, they still suffer from forgetting when incrementally learning a sequence of tasks. To alleviate this problem, recent works enhance existing models by sparse experience replay and local adaption, which yield satisfactory performance. However, in this paper we find that pre-trained language models like BERT have a potential ability to learn sequentially, even without any sparse memory replay. To verify the ability of BERT to maintain old knowledge, we adopt and re-finetune single-layer probe networks with the parameters of BERT fixed. We investigate the models on two types of NLP tasks, text classification and extractive question answering. Our experiments reveal that BERT can actually generate high quality representations for previously learned tasks in a long term, under extremely sparse replay or even no replay. We further introduce a series of novel methods to interpret the mechanism of forgetting and how memory rehearsal plays a significant role in task incremental learning, which bridges the gap between our new discovery and previous studies about catastrophic forgetting 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Demystifying Language Model Forgetting with Low-rank Example AssociationsXisen Jin, Xiang RenNeurIPS 2025 · 被引用 9 次
- What Will My Model Forget? Forecasting Forgotten Examples in Language Model RefinementXisen Jin, Xiang RenICML 2024 · 被引用 8 次
- Spurious Forgetting in Continual Learning of Language ModelsJunhao Zheng, Xidi Cai, Shengjie Qiu, Qianli MaICLR 2025
- Learn or Recall? Revisiting Incremental Learning with Pre-trained Language ModelsJunhao Zheng, Shengjie Qiu, Qianli MaACL 2024
- MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMsHe Li, Haoang Chi, Qizhou Wang, Yunxin Mao 等ICML 2026
它引用的顶会 Paper6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Improving Neural Language Generation with Spectrum ControlLingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu 等ICLR 2020 · 被引用 94 次
- Pretrained Language Model in Continual Learning: A Comparative StudyTongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li 等ICLR 2022 · 被引用 76 次
- Efficient Meta Lifelong-Learning with Limited MemoryZirui Wang, Sanket Vaibhav Mehta, Barnabás Póczos, Jaime G. CarbonellEMNLP 2020 · 被引用 47 次
相关 Paper
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual LearningHuihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu 等ICML 2026 · 被引用 14 次
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen 等AAAI 2023 · 被引用 4 次
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 被引用 212 次
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal 等NeurIPS 2022 · 被引用 75 次
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che 等EMNLP 2020 · 被引用 152 次
