Spurious Forgetting in Continual Learning of Language Models
Junhao Zheng, Xidi Cai, Shengjie Qiu, Qianli Ma
Abstract
Recent advancements in large language models (LLMs) reveal a perplexing phenomenon in continual learning: despite extensive training, models experience significant performance declines, raising questions about task alignment and underlying knowledge retention. This study first explores the concept of "spurious forgetting", proposing that such performance drops often reflect a decline in task alignment rather than knowledge loss. Through controlled experiments with a synthesized dataset, we investigate the dynamics of model performance during the initial training phases of new tasks, discovering that early optimization steps can disrupt previously established task alignments. Our theoretical analysis connects these shifts to orthogonal updates in model weights, providing a robust framework for understanding this behavior. Ultimately, we introduce a Freezing strategy that fix the bottom layers of the model, leading to substantial improvements in four continual learning scenarios. Our findings underscore the critical distinction between task alignment and knowledge retention, paving the way for more effective strategies in continual learning. The source code is publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33a0074d-d6c2-4fb7-8a0e-8f08234fbbb7Cited by top-tier papers18
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu et al.ICML 2026 · 102 citations
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMsZhixin Xie, Xurui Song, Jun LuoNeurIPS 2025 · 11 citations
- Demystifying Language Model Forgetting with Low-rank Example AssociationsXisen Jin, Xiang RenNeurIPS 2025 · 9 citations
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action ModelZhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen et al.EMNLP 2025 · 6 citations
- Can Large Language Models Master Complex Card Games?Wei Wang, Fuqing Bie, Junzhe Chen, Dan Zhang et al.NeurIPS 2025 · 6 citations
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsDingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen et al.EMNLP 2025 · 4 citations
- Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual LearningYitong Wang, Xue Han, Wenchun Gao, Qian Hu et al.ACL 2026
- Merge before Forget: A Single LoRA Continual Learning via Continual MergingFuli Qiao, Mehrdad MahdaviICLR 2026 · 11 citations
- Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual LearningNikhil Shivakumar Nayak, Krishnateja Killamsetty, Ligong Han, Abhishek Bhandwaldar et al.ICLR 2026 · 17 citations
- SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual LearningLina Yang, Yusheng Liao, Yanfeng Wang, Yu WangACL 2026
