Targeted Low-rank Refinement: Enhancing Sparse Language Models with Precision
Li Shen, Anke Tang, Yong Luo, Tao Sun, Han Hu, Xiaochun Cao
Abstract
Pruning is a widely used technique for compressing large neural networks that eliminates weights with minimal impact on performance. Current pruning methods, exemplified by magnitude pruning, assign importance scores to weights based on their magnitude and remove those below a certain threshold. However, these methods introduce a gap between the original dense and pruned sparse models, potentially impairing performance, especially at high sparsity ratios. To address this issue, we introduce a method that bridges this gap through low-rank approximation of the difference between dense and sparse matrices. Our approach iteratively refines the sparse weight matrix with a low-rank adjustment, capturing essential information typically lost during pruning. We provide a comprehensive theoretical analysis of our method, establishing its convergence properties and efficacy. Experimental results on LLaMA models validate our method's effectiveness across various pruning techniques and sparsity levels. At 50% sparsity, it reduces perplexity by 53.9% compared to conventional magnitude pruning on LLaMA-7B. Furthermore, our approach enables an 8.6% reduction in model parameters while maintaining a sparsity ratio of about 50%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ebeb27e-0c89-4bb8-b752-97ddced4a844Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
Related papers
- Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude CompensationXinrui Chen, Hongxing Zhang, Fanyi Zeng, Yongxian Wei et al.AAAI 2026 · 3 citations
- Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse FactorizationVladimír Boza, Vladimír MackoICLR 2025
- DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance ConcentrationTianteng Gu, Bei Liu, Bo Xiao, Ke Zeng et al.NeurIPS 2025 · 7 citations
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- Compress Large Language Models via Collaboration Between Learning and Matrix ApproximationYuesen Liao, Zhiwei Li, Binrui Wu, Zihao Cheng et al.NeurIPS 2025 · 1 citation
