Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language Models
Kai Tong, Kang Pan, Xiao Zhang, Erli Meng, Run He, Yawen Cui, Nuoyan Guo, Huiping Zhuang
Abstract
Fine-tuning large language models (LLMs) on data of various domains encompassing advanced language-related tasks. Such a process is essentially a continual learning procedure since it is impractical to re-access the pre-trained data in most parties. This naturally attracts catastrophic forgetting of LLMs' previously learned knowledge, such as general skills. Existing techniques either leverage previous data to replay, leading to extra computational costs, or utilize a single parameter-efficient module to learn the downstream task, constraining new knowledge absorption with interference among different tasks. To handle these issues, we propose an Analytic Subspace Routing (Any-SSR) to address these challenges. For each task, we isolate the learning within a subspace of deep layers' features via low-rank adaptation, eliminating knowledge interference between different tasks. Additionally, we propose an analytic routing mechanism to properly utilize knowledge learned in different subspaces. Our approach employs Recursive Least Squares to train a multi-task router model, allowing the router to dynamically adapt to incoming data without requiring access to historical data. Subsequently, the router effectively assigns the current task to an appropriate subspace and has a non-forgetting property of previously learned tasks with a solid theoretical guarantee. Experimental results demonstrate that our method achieves near-perfect retention of prior knowledge while seamlessly integrating new information, effectively mitigating the core limitations of existing methods. Our code is available at https://github.com/ZHUANGHP/Any-SSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5Chengwei Qin, Shafiq R. JotyICLR 2022 · 128 citations
- ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy ProtectionHuiping Zhuang, Zhenyu Weng, Hongxin Wei, Renchunzi Xie et al.NeurIPS 2022 · 106 citations
Related papers
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual LearningYitong Wang, Xue Han, Wenchun Gao, Qian Hu et al.ACL 2026
- Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language ModelsYuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang et al.ACL 2025 · 7 citations
- KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction TuningLingyun Song, Ziyao Chen, Kang Pan, Xiaolin Han et al.AAAI 2026
- SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual LearningLina Yang, Yusheng Liao, Yanfeng Wang, Yu WangACL 2026
