ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
Xiang Meng, Kayhan Behdin, Haoyue Wang, Rahul Mazumder
Abstract
The impressive performance of Large Language Models (LLMs) across various natural language processing tasks comes at the cost of vast computational resources and storage requirements. One-shot pruning techniques offer a way to alleviate these burdens by removing redundant weights without the need for retraining. Yet, the massive scale of LLMs often forces current pruning approaches to rely on heuristics instead of optimization-based techniques, potentially resulting in suboptimal compression. In this paper, we introduce ALPS, an optimization-based framework that tackles the pruning problem using the operator splitting technique and a preconditioned conjugate gradient-based post-processing step. Our approach incorporates novel techniques to accelerate and theoretically guarantee convergence while leveraging vectorization and GPU parallelism for efficiency. ALPS substantially outperforms state-of-the-art methods in terms of the pruning objective and perplexity reduction, particularly for highly sparse models. On the OPT-30B model with 70% sparsity, ALPS achieves a 13% reduction in test perplexity on the WikiText dataset and a 19% improvement in zero-shot benchmark performance compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 042b611a-8fef-4f17-b43d-bf33d725db66Cited by top-tier papers20
- Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model AlignmentZhipeng Chen, Kun Zhou, Xin Zhao, Jingyuan Wang et al.EMNLP 2024 · 7 citations
- Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution OptimizationGuanchen Li, Yixing Xu, Zeping Li, Ji Liu et al.NeurIPS 2025 · 7 citations
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-ShotKaiwen TUO, Huan WangICML 2026 · 6 citations
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMMKwanhee Lee, Hyeondo Jang, Dongyeop Lee, Dan Alistarh et al.ICLR 2026 · 5 citations
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought ReconstructionRyan Lucas, Kayhan Behdin, Zhipeng Wang, Qingquan Song et al.ICLR 2026 · 3 citations
Builds on23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- WRP: Weight Recover Prune for Structured SparsityZhendong Tan, Xingjun Zhang, Zheng WeiACL 2024
- A Robust Optimization Guided Pruning Framework for Vision and Large Language ModelsGabriel Afriat, Hussein Hazimeh, Dimitris Paparas, Rahul MazumderICML 2026
- Multi-Objective One-Shot Pruning for Large Language ModelsWeiyu Chen, Hansi Yang, Yunhao Gou, Han Shi et al.NeurIPS 2025
- OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial OptimizationXiang Meng, Shibal Ibrahim, Kayhan Behdin, Hussein Hazimeh et al.ICML 2024 · 17 citations
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
