Pretrained Language Model in Continual Learning: A Comparative Study
Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, Gholamreza Haffari
Abstract
Continual learning (CL) is a setting in which a model learns from a stream of incoming data while avoiding to forget previously learned knowledge. Pre-trained language models (PLMs) have been successfully employed in continual learning of different natural language problems. With the rapid development of many continual learning methods and PLMs, understanding and disentangling their interactions become essential for continued improvement of continual learning performance. In this paper, we thoroughly compare the continual learning performance over the combination of 5 PLMs and 4 CL approaches on 3 benchmarks in 2 typical incremental settings. Our extensive experimental analyses reveal interesting performance differences across PLMs and across CL methods. Furthermore, our representativeness probing analyses dissect PLMs’ performance characteristics in a layer-wise and task-wise manner, uncovering the extent to which their inner layers suffer from forgetting, and the effect of different CL approaches on each layer. Finally, our observations and analyses open up a number of important research questions that will inform and guide the design of effective continual learning techniques.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1396fde9-0c06-4ff2-bfdc-747464b24ad3Cited by top-tier papers24
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang et al.NeurIPS 2024 · 47 citations
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang et al.ICML 2023 · 21 citations
- Mitigating the Alignment Tax of RLHFYong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao et al.EMNLP 2024 · 18 citations
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual LearningHuihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu et al.ICML 2026 · 14 citations
Related papers
- Enhancing Visual Continual Learning with Language-Guided SupervisionBolin Ni, Hongbo Zhao, Chenghao Zhang, Ke Hu et al.CVPR 2024 · 8 citations
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa et al.ICLR 2023 · 15 citations
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model MergingBowen Wang, Haiyuan Wan, Liwen Shi, Chen Yang et al.EMNLP 2025
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding et al.AAAI 2026
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
