Targeted Low-rank Refinement: Enhancing Sparse Language Models with Precision
Li Shen, Anke Tang, Yong Luo, Tao Sun, Han Hu, Xiaochun Cao
摘要
Pruning is a widely used technique for compressing large neural networks that eliminates weights with minimal impact on performance. Current pruning methods, exemplified by magnitude pruning, assign importance scores to weights based on their magnitude and remove those below a certain threshold. However, these methods introduce a gap between the original dense and pruned sparse models, potentially impairing performance, especially at high sparsity ratios. To address this issue, we introduce a method that bridges this gap through low-rank approximation of the difference between dense and sparse matrices. Our approach iteratively refines the sparse weight matrix with a low-rank adjustment, capturing essential information typically lost during pruning. We provide a comprehensive theoretical analysis of our method, establishing its convergence properties and efficacy. Experimental results on LLaMA models validate our method's effectiveness across various pruning techniques and sparsity levels. At 50% sparsity, it reduces perplexity by 53.9% compared to conventional magnitude pruning on LLaMA-7B. Furthermore, our approach enables an 8.6% reduction in model parameters while maintaining a sparsity ratio of about 50%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
相关 Paper
- Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude CompensationXinrui Chen, Hongxing Zhang, Fanyi Zeng, Yongxian Wei 等AAAI 2026 · 被引用 3 次
- Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse FactorizationVladimír Boza, Vladimír MackoICLR 2025
- DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance ConcentrationTianteng Gu, Bei Liu, Bo Xiao, Ke Zeng 等NeurIPS 2025 · 被引用 7 次
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang 等ICML 2025
- Compress Large Language Models via Collaboration Between Learning and Matrix ApproximationYuesen Liao, Zhiwei Li, Binrui Wu, Zihao Cheng 等NeurIPS 2025 · 被引用 1 次
