Merge before Forget: A Single LoRA Continual Learning via Continual Merging
Fuli Qiao, Mehrdad Mahdavi
Abstract
Parameter-efficient continual learning has emerged as a promising approach for large language models (LLMs) to mitigate catastrophic forgetting while enabling adaptation to new tasks. Current Low-Rank Adaptation (LoRA) continual learning techniques often retain and freeze previously learned LoRAs or generate data representations to overcome forgetting, typically utilizing these to support new LoRAs learn new tasks. However, these methods not only ignore growing computational memory with tasks and limited storage space but also suffer from potential task interference due to the lack of effective LoRA merging mechanisms. In this paper, we propose a novel continual learning method that orthogonally initializes and sequentially merges LoRAs updates into a single unified LoRA. Our method leverages orthogonal basis extraction from previously learned LoRA to initialize the learning of new tasks, further exploits the intrinsic asymmetry property of LoRA components by using a time-aware scaling mechanism to balance new and old knowledge during continual merging. Our approach maintains constant memory complexity with respect to the number of tasks, minimizes interference between past and new tasks via orthogonal basis initialization, and improves performance over asymmetric LoRA merging via adaptive scaling. We provide theoretical analysis to justify our design and conduct extensive experiments across diverse continual learning benchmarks using various Llama models, demonstrating the effectiveness and efficiency of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on31
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
Related papers
- Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual LearningYitong Wang, Xue Han, Wenchun Gao, Qian Hu et al.ACL 2026
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language ModelsYan-Shuo Liang, Jia-Rui Chen, Wu-Jun LiNeurIPS 2025 · 15 citations
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual LearningHaomiao Qiu, Miao Zhang, Ziyue Qiao, Liqiang NieNeurIPS 2025 · 8 citations
- Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language ModelsYuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang et al.ACL 2025 · 7 citations
