LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
Zihang Liu, Tianyu Pang, Oleg Balabanov, Chaoqun Yang, Tianjin Huang, Lu Yin, Yaoqing Yang, Shiwei Liu
摘要
Recent studies have shown that supervised finetuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabilities. However, full fine-tuning (Full FT), while powerful, is computationally expensive and susceptible to overfitting and catastrophic forgetting, particularly when data is limited. Sparse fine-tuning, which previously achieved notable success by updating only a small subset of model parameters, offers a promising trade-off between efficiency and effectiveness. Yet, it has lagged behind in the LLM era due to the difficulty of identifying parameters truly critical for reasoning. In this work, we state that weights with the largest magnitude after low-rank approximation are critical weights for fine-tuning, which we call Principal Weights. Surprisingly, while magnitude-based sparse finetuning performs poorly as a baseline on LLM fine-tuning, it becomes highly effective after rank reduction. These insights motivate our method: Low-rank Informed Sparse Fine-Tuning (LIFT). LIFT only updates the top 5% Principal Weights throughout training and consistently achieves better performance on reasoning tasks than Full FT, while maintaining memory efficiency on par with popular parameter-efficient fine-tuning methods. In addition to strong performance on target domains such as arithmetic reasoning, LIFT also retains up to 20% more source-domain knowledge, compared to Full FT and LoRA. Our code is available at: https://github.com/zihanghliu/LIFT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning ModelsNan Zhang, Eugene Kwek, Yusen Zhang, Hieu Nguyen 等ICLR 2026 · 被引用 5 次
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language ModelsYuxuan Gu, Wuyang Zhou, Giorgos Iacovides, Danilo P. MandicACL 2026 · 被引用 2 次
- GeoRA: Geometry-Aware Low-Rank Adaptation for RLVRJiaying Zhang, Lei Shi, Jiguo Li, Jun Xu 等ACL 2026 · 被引用 1 次
- Eigenspectrum Analysis of Neural Networks without Aspect Ratio BiasYuanzhe Hu, Kinshuk Goel, Vlad Killiakov, Yaoqing YangICML 2025
它引用的顶会 Paper31
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsFanxu Meng, Zhaohui Wang, Muhan ZhangNeurIPS 2024 · 被引用 374 次
- Training Neural Networks with Fixed Sparse MasksYi-Lin Sung, Varun Nair, Colin RaffelNeurIPS 2021 · 被引用 295 次
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
相关 Paper
- S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured SparsityXinyu Yang, Jixuan Leng, Geyang Guo, Jiawei Zhao 等NeurIPS 2024 · 被引用 13 次
- Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuningChenxi Huang, Shaotian Yan, Liang Xie, Binbin Lin 等ACL 2025
- RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationMahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan AlistarhICML 2024 · 被引用 53 次
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual SparsitySamir Khaki, Xiuyu Li, Junxian Guo, Ligeng Zhu 等ICML 2025
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
