Spurious Forgetting in Continual Learning of Language Models
Junhao Zheng, Xidi Cai, Shengjie Qiu, Qianli Ma
摘要
Recent advancements in large language models (LLMs) reveal a perplexing phenomenon in continual learning: despite extensive training, models experience significant performance declines, raising questions about task alignment and underlying knowledge retention. This study first explores the concept of "spurious forgetting", proposing that such performance drops often reflect a decline in task alignment rather than knowledge loss. Through controlled experiments with a synthesized dataset, we investigate the dynamics of model performance during the initial training phases of new tasks, discovering that early optimization steps can disrupt previously established task alignments. Our theoretical analysis connects these shifts to orthogonal updates in model weights, providing a robust framework for understanding this behavior. Ultimately, we introduce a Freezing strategy that fix the bottom layers of the model, leading to substantial improvements in four continual learning scenarios. Our findings underscore the critical distinction between task alignment and knowledge retention, paving the way for more effective strategies in continual learning. The source code is publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu 等ICML 2026 · 被引用 102 次
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMsZhixin Xie, Xurui Song, Jun LuoNeurIPS 2025 · 被引用 11 次
- Demystifying Language Model Forgetting with Low-rank Example AssociationsXisen Jin, Xiang RenNeurIPS 2025 · 被引用 9 次
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action ModelZhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen 等EMNLP 2025 · 被引用 6 次
- Can Large Language Models Master Complex Card Games?Wei Wang, Fuqing Bie, Junzhe Chen, Dan Zhang 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsDingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen 等EMNLP 2025 · 被引用 4 次
- Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual LearningYitong Wang, Xue Han, Wenchun Gao, Qian Hu 等ACL 2026
- Merge before Forget: A Single LoRA Continual Learning via Continual MergingFuli Qiao, Mehrdad MahdaviICLR 2026 · 被引用 11 次
- Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual LearningNikhil Shivakumar Nayak, Krishnateja Killamsetty, Ligong Han, Abhishek Bhandwaldar 等ICLR 2026 · 被引用 17 次
- SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual LearningLina Yang, Yusheng Liao, Yanfeng Wang, Yu WangACL 2026
