SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
Xudong Lu, Aojun Zhou, Yuhui Xu, Renrui Zhang, Peng Gao, Hongsheng Li
Abstract
Large Language Models (LLMs) have become pivotal in advancing the field of artificial intelligence, yet their immense sizes pose significant challenges for both fine-tuning and deployment. Current post-training pruning methods, while reducing the sizes of LLMs, often fail to maintain their original performance. To address these challenges, this paper introduces SPP, a Sparsity-Preserved Parameter-efficient fine-tuning method. Different from existing post-training pruning approaches that struggle with performance retention, SPP proposes to employ lightweight learnable column and row matrices to optimize sparse LLM weights, keeping the structure and sparsity of pruned pre-trained models intact. By elementwise multiplication and residual addition, SPP ensures the consistency of model sparsity pattern and ratio during both training and weight-merging processes. We demonstrate the effectiveness of SPP by applying it to the LLaMA and LLaMA-2 model families with recent post-training pruning methods. Our results show that SPP significantly enhances the performance of models with different sparsity patterns (i.e. unstructured and N:M sparsity), especially for those with high sparsity ratios (e.g. 75%), making it a promising solution for the efficient fine-tuning of sparse LLMs. Code will be made available at https: //github.com/Lucky-Lance/SPP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e647ec3-86cb-4ab0-8ae2-52965d953f0fCited by top-tier papers12
- MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling QuantizationAkshat Ramachandran, Souvik Kundu, Tushar KrishnaISCA 2025 · 10 citations
- Entropy-Based Block Pruning for Efficient Large Language ModelsLiangwei Yang, Yuhui Xu, Juntao Tan, Doyen Sahoo et al.ICLR 2026 · 3 citations
- Mitigating Non-IID Drift in Zeroth-Order Federated LLM Fine-Tuning with Transferable SparsityYide Ran, Wentao Guo, Jingwei Sun, Yanzhou Pan et al.ICLR 2026 · 1 citation
- BaWA: Automatic Optimizing Pruning Metric for Large Language Models with Balanced Weight and ActivationLian Liu, Xiandong Zhao, Guanchen Li, Dong Li et al.ICML 2025
- Effective Interplay between Sparsity and Quantization: From Theory to PracticeSimla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin et al.ICLR 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
Related papers
- PAT: Pruning-Aware Tuning for Large Language ModelsYijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang et al.AAAI 2025 · 1 citation
- APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and InferenceBowen Zhao, Hannaneh Hajishirzi, Qingqing CaoICML 2024 · 31 citations
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang et al.ICLR 2026 · 24 citations
- GPTailor: Large Language Model Pruning Through Layer Cutting and StitchingGuinan Su, Li Shen, Lu Yin, Shiwei Liu et al.ICLR 2026 · 3 citations
