Pruning Large Language Models with Semi-Structural Adaptive Sparse Training
Weiyu Huang, Yuezhou Hu, Guohao Jian, Jun Zhu, Jianfei Chen
Abstract
The remarkable success of Large Language Models (LLMs) relies heavily on their substantial scale, which poses significant challenges during model deployment in terms of latency and memory consumption. Recently, numerous studies have attempted to compress LLMs using one-shot pruning methods. However, these methods often suffer from considerable performance degradation on complex language understanding tasks, raising concerns about the feasibility of pruning in LLMs. To address this issue, we propose Adaptive Sparse Trainer (AST), a novel and efficient retraining framework tailored for semi-structured sparse models. AST enables models to learn optimal masks during the weight update process without incurring additional computational overhead. Furthermore, we demonstrate that incorporating knowledge distillation significantly improves retraining efficiency and enhances model performance under fixed computational constraints. Additionally, a supplementary set of well-initialized parameters is integrated to further augment the model's efficacy. AST achieves state-of-the-art performance with minimal training cost. When applied to the LLaMA2-7B model, AST reduces the perplexity and zero-shot accuracy gap between dense and 2:4 semi-structured sparse models to 0.6 and 1.16%, respectively, utilizing less than 0.4% of the pretraining tokens and GPU hours. Our work demonstrates the feasibility of deploying semi-structured sparse LLMs and offers a promising alternative for achieving highly compressed models when combined with existing quantization techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22b1e593-0d50-4dfd-94c5-38c3ae954735Cited by top-tier papers8
- Mitigating Non-IID Drift in Zeroth-Order Federated LLM Fine-Tuning with Transferable SparsityYide Ran, Wentao Guo, Jingwei Sun, Yanzhou Pan et al.ICLR 2026 · 1 citation
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMsYan Sun, Qixin Zhang, Zhiyuan Yu, Xikun Zhang et al.ICLR 2026 · 1 citation
- Deterministic Differentiable Structured Pruning for Large Language ModelsWeiyu Huang, Pengle Zhang, Xiaolu Zhang, JUN ZHOU et al.ICML 2026 · 1 citation
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion ModelsXing Cong, Hanlin Tang, Kan Liu, Lan Tao et al.ICML 2026 · 1 citation
- Proxsparse: Regularized Learning of Semi-Structured Sparsity masks for Pretrained LLMSHongyi Liu, Rajarshi Saha, Zhen Jia, Youngsuk Park et al.ICML 2025
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
Related papers
- LLaMaFlex: Many-in-one LLMs via Generalized Pruning and Weight SharingRuisi Cai, Saurav Muralidharan, Hongxu Yin, Zhangyang Wang et al.ICLR 2025
- FISTAPruner: Layer-wise Post-training Pruning for Large Language ModelsPengxiang Zhao, Hanyu Hu, Ping Li, Yi Zheng et al.EMNLP 2025
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski et al.NeurIPS 2024 · 198 citations
