Fluctuation-Based Adaptive Structured Pruning for Large Language Models
Yongqi An, Xu Zhao, Tao Yu, Ming Tang, Jinqiao Wang
Abstract
Network Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost all of the existing retraining-free pruning approaches for LLMs focus on unstructured pruning, which requires specific hardware support for acceleration. In this paper, we propose a novel retraining-free structured pruning framework for LLMs, named FLAP (FLuctuation-based Adaptive Structured Pruning). It is hardware-friendly by effectively reducing storage and enhancing inference speed. For effective structured pruning of LLMs, we highlight three critical elements that demand the utmost attention: formulating structured importance metrics, adaptively searching the global compressed model, and implementing compensation mechanisms to mitigate performance loss. First, FLAP determines whether the output feature map is easily recoverable when a column of weight is removed, based on the fluctuation pruning metric. Then it standardizes the importance scores to adaptively determine the global compressed model structure. At last, FLAP adds additional bias terms to recover the output feature maps using the baseline values. We thoroughly evaluate our approach on a variety of language benchmarks. Without any retraining, our method significantly outperforms the state-of-the-art methods, including LLM-Pruner and the extension of Wanda in structured pruning. The code is released at https://github.com/CASIA-IVA-Lab/FLAP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c2eeacf-043d-4539-b04d-4d65da88222eCited by top-tier papers67
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu et al.NeurIPS 2024 · 51 citations
- Search for Efficient Large Language ModelsXuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong et al.NeurIPS 2024 · 23 citations
- Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance AssessmentJun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang et al.AAAI 2025 · 23 citations
- AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant DeploymentYonggan Fu, Zhongzhi Yu, Junwei Li, Jiayi Qian et al.NeurIPS 2024 · 15 citations
- SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model CompressionXinhao Huang, You-Liang Huang, Zeyi WenAAAI 2025 · 14 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
Related papers
- Structured Optimal Brain Pruning for Large Language ModelsJiateng Wei, Quan Lu, Ning Jiang, Siqi Li et al.EMNLP 2024 · 2 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- SlimLLM: Accurate Structured Pruning for Large Language ModelsJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2025
- SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language ModelsXun Liang, Hanyu Wang, Huayi Lai, Simin Niu et al.AAAI 2026
- Olica: Efficient Structured Pruning of Large Language Models without RetrainingJiujun He, Huazhen LinICML 2025
