Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy Gradient
Yuan Gao, Zujing Liu, Weizhong Zhang, Bo Du, Gui-Song Xia
Abstract
Recent Large-Language Models (LLMs) pruning methods typically operate at the posttraining phase without the expensive weight finetuning, however, their pruning criteria often rely on heuristically hand-crafted metrics, potentially leading to suboptimal performance. We instead propose a novel optimizationbased structural pruning that learns the pruning masks in a probabilistic space directly by optimizing the loss of the pruned model. To preserve efficiency, our method eliminates the back-propagation through the LLM per se during optimization, requiring only the forward pass of the LLM. We achieve this by learning an underlying Bernoulli distribution to sample binary pruning masks, where we decouple the Bernoulli parameters from LLM loss, facilitating efficient optimization via policy gradient estimator without back-propagation. Thus, our method can 1) support global and heterogeneous pruning (i.e., automatically determine different redundancy for different layers), and 2) optionally initialize with a metric-based method (for our Bernoulli distributions). Extensive experiments conducted on LLaMA, LLaMA-2, LLaMA-3, Vicuna, and Mistral models using the C4 and WikiText2 datasets demonstrate the promising performance of our method in efficiency and effectiveness. Code is available at https://github.com/ ethanygao/backprop-free_LLM_pruning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMsYan Sun, Qixin Zhang, Zhiyuan Yu, Xikun Zhang et al.ICLR 2026 · 1 citation
- Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain AdaptationTzu Ling Liu, Ian Stavness, Mrigank RochanCVPR 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
Related papers
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language ModelsMingge Lu, Jingwei Sun, Junqing Lin, Zechun Zhou et al.NeurIPS 2025 · 1 citation
- Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution OptimizationGuanchen Li, Yixing Xu, Zeping Li, Ji Liu et al.NeurIPS 2025 · 7 citations
- Dual-Assessment Driven Pruning: Iterative Optimizing Layer-wise Sparsity for Large Language ModelQinghui Sun, Weilun Wang, Yanni Zhu, Shenghuan He et al.KDD 2024 · 3 citations
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu et al.ICML 2024 · 64 citations
- Structured Optimal Brain Pruning for Large Language ModelsJiateng Wei, Quan Lu, Ning Jiang, Siqi Li et al.EMNLP 2024 · 2 citations
