Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMs
Jun Wang, Liang Ding, Shuai Wang, Hongyu Li, Yong Luo, Huangxuan Zhao, Han Hu, Bo Du
Abstract
Continual learning for large language models (LLMs) demands a precise balance between plasticity -the ability to absorb new tasks -and stability -the preservation of previously learned knowledge. Conventional rehearsal methods, which replay stored examples, are limited by long-term data inaccessibility; earlier pseudorehearsal methods require additional generation modules, while self-synthesis approaches often generate samples that poorly align with real tasks, suffer from unstable outputs, and ignore task relationships. We present Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity (SERS), a lightweight framework that 1) decouples pseudo-input synthesis from label creation, using semantic masking and template guidance to produce diverse, task-relevant prompts without extra modules; 2) applies label self-evolution, blending base-model priors with fine-tuned outputs to prevent over-specialization; and 3) introduces a dynamic regularizer driven by the Wasserstein distance between task distributions, automatically relaxing or strengthening constraints in proportion to task similarity. Experiments across diverse tasks on different LLMs show that our SERS reduces forgetting by over 2% points against strong pseudo-rehearsal baselines, by ensuring efficient data utilization and wisely transferring knowledge. The code will be released at https://github.com/JerryWangJun/LLM_CL_SERS/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d3eba4b-16d4-41b9-8f3b-818b962a26cfCited by top-tier papers1
Ask how each one uses itBuilds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 267 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi et al.EMNLP 2022 · 238 citations
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik et al.ICLR 2024 · 45 citations
Related papers
- Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang et al.ACL 2024 · 13 citations
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa et al.ICLR 2023 · 15 citations
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao et al.EMNLP 2024 · 4 citations
- Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMsHuanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang et al.ACL 2026
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model MergingBowen Wang, Haiyuan Wan, Liwen Shi, Chen Yang et al.EMNLP 2025
