Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models
Yingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao, Lu Hou, Carlo Vittorio Cannistraci
Abstract
With the rapid growth of large language models (LLMs), there is increasing demand for memory and computation in LLMs. Recent efforts on post-training pruning of LLMs aim to reduce the model size and computation requirements, yet the performance is still sub-optimal. In this paper, we present a plug-and-play solution for post-training pruning of LLMs. The proposed solution has two innovative components: 1) Relative Importance and Activations (RIA), a new pruning metric that jointly considers the weight and activations efficiently on LLMs; and 2) Channel Permutation, a new approach to maximally preserve important weights under N:M sparsity. The two proposed components can be readily combined to further enhance the N:M semi-structured pruning of LLMs. Our empirical experiments show that RIA alone can already surpass all existing post-training pruning methods on prevalent LLMs, e.g., LLaMA ranging from 7B to 65B. Furthermore, N:M semi-structured pruning with channel permutation can even outperform the original LLaMA2-70B on zero-shot tasks, together with practical speedup on specific hardware. Our code is available at: https://github.com/biomedicalcybernetics/Relative-importance-and-activation-pruning * This work is partially done during the internship at Huawei Noah's Ark Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers43
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich et al.NeurIPS 2024 · 72 citations
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu et al.ICML 2024 · 64 citations
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu et al.NeurIPS 2024 · 51 citations
- Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance AssessmentJun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang et al.AAAI 2025 · 23 citations
Builds on16
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
Related papers
- PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsLancheng Zou, Shuo Yin, Zehua Pei, Tsung-Yi Ho et al.NeurIPS 2025 · 1 citation
- SlimLLM: Accurate Structured Pruning for Large Language ModelsJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2025
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer ConcatenationFei Wang, Li Shen, Liang Ding, Chao Xue et al.NeurIPS 2025 · 7 citations
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang et al.ICLR 2026 · 24 citations
- Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang et al.AAAI 2024 · 130 citations
