LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction Error
Zhiguo Yang, Changjian Deng, Qinke Chen, Zijing Zhou, Jian Cheng
摘要
Deploying large language models (LLMs) on platforms with insufficient computational resources remains a key challenge. Weight pruning is an efficient model compression technique that can reduce model size without retraining LLMs. However, due to the massive number of parameters, it is infeasible to estimate the importance of weights globally, and most prior studies assign a uniform sparsity ratio across all layers. Recent findings reveal that layers contribute unevenly to LLM performance, making it necessary to investigate layer-wise importance. Existing layer-wise sparsity allocation methods, such as OWL and DLP, rely on weight scoring and carefully designed score proxies to estimate layer-wise importance and sparsity ratios, while enforcing identical sparsity to blocks and projection weights within a layer to avoid performance degradation. In this work, we propose layer-wise Sparsity Allocation (LSA) for LLM pruning, which quantifies layer-wise importance by evaluating the minimal linear reconstruction error of each transformer layer under the assumption that 50% of its least important weights are removed. Moreover, our method supports non-uniform sparsity allocation at block-or projection-level granularity within layers, without incurring catastrophic performance degradation. Experimental results demonstrate that LSA maintains high performance at high sparsity levels. At an overall sparsity ratio of 70%, LSA surpasses state-of-the-art methods across language modeling tasks and seven zero-shot tasks. Code is available at https://github.com/BeiYazi0/LSA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
相关 Paper
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu 等NeurIPS 2024 · 被引用 51 次
- BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity AllocationPeng Xu, Wenqi Shao, Mengzhao Chen, Shitao Tang 等ICLR 2024 · 被引用 52 次
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang 等ICML 2025
- Adaptive Layer Sparsity for Large Language Models via Activation Correlation AssessmentWei Li, Lujun Li, Mark G. Lee, Shengjie SunNeurIPS 2024 · 被引用 39 次
- Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High SparsityLu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh 等ICML 2024 · 被引用 183 次
