Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?
Hyowon Wi, Noseong Park
Abstract
In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. In this work, we uncover several key insights regarding the singular components of network parameters based on Singular Value Decomposition (SVD). Firstly, the principal singular components with large singular values in pre-trained network parameters can be effectively reused during finetuning, whereas the minor components with smaller singular values are more task-specific and require substantial adaptation. Secondly, we first establish the theoretical connection that the uncontrolled growth of singular values in LoRA adapters leads to the forgetting of pretrained knowledge -a well-known issue referred to as catastrophic forgetting. Building on these observations, we propose SCLoRA, which injects parameterized singular components with spectral clipping into the pre-trained model in a way that is aware of the spectral distribution of the pre-trained model. SCLoRA effectively adapts to new tasks by focusing updates on components that require adaptation, while simultaneously alleviating catastrophic forgetting. We conduct extensive experiments and demonstrate that SCLoRA not only improves downstream performance but also effectively retains pre-trained knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b732b4e9-a278-4d69-b058-af251698d1daBuilds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting During Parameter-Efficient Fine-TuningYifeng Xiong, Xiaohui XieAAAI 2026 · 6 citations
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 152 citations
- L2-LoRA: Improving Low-Rank Adaptation with Layer-Specific RegularizationXiang Zhang, Rui Xie, Shikun ZhangAAAI 2026
- Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained KnowledgePengwei Tang, Xiaolin Hu, Yong Liu, Lizhong Ding et al.AAAI 2026 · 6 citations
- Efficient Fine-Tuning of Large Models Via Nested Low-Rank AdaptationLujun Li, Cheng Lin, Dezhi Li, You-Liang Huang et al.ICCV 2025 · 1 citation
