Fast as CHITA: Neural Network Pruning with Combinatorial Optimization
Riade Benbaki, Wenyu Chen, Xiang Meng, Hussein Hazimeh, Natalia Ponomareva, Zhe Zhao, Rahul Mazumder
Abstract
The sheer size of modern neural networks makes model serving a serious computational challenge. A popular class of compression techniques overcomes this challenge by pruning or sparsifying the weights of pretrained networks. While useful, these techniques often face serious tradeoffs between computational requirements and compression quality. In this work, we propose a novel optimization-based pruning framework that considers the combined effect of pruning (and updating) multiple weights subject to a sparsity constraint. Our approach, CHITA, extends the classical Optimal Brain Surgeon framework and results in significant improvements in speed, memory, and performance over existing optimization-based approaches for network pruning. CHITA's main workhorse performs combinatorial optimization updates on a memory-friendly representation of local quadratic approximation(s) of the loss function. On a standard benchmark of pretrained models and datasets, CHITA leads to significantly better sparsity-accuracy tradeoffs than competing methods. For example, for MLPNet with only 2% of the weights retained, our approach improves the accuracy by 63% relative to the state of the art. Furthermore, when used in conjunction with fine-tuning SGD steps, our method achieves significant accuracy gains over the state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b46defb-2843-476c-8ae4-5b5124b836bdCited by top-tier papers20
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh et al.ICLR 2024 · 48 citations
- ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language ModelsXiang Meng, Kayhan Behdin, Haoyue Wang, Rahul MazumderNeurIPS 2024 · 19 citations
- How Sparse Can We Prune A Deep Network: A Fundamental Limit PerspectiveQiaozhe Zhang, Ruijie Zhang, Jun Sun, Yingzhuang LiuNeurIPS 2024 · 14 citations
- The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order InformationDiyuan Wu, Ionut-Vlad Modoranu, Mher Safaryan, Denis Kuznedelev et al.NeurIPS 2024 · 8 citations
Builds on3
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
- The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural NetworksXin Yu, Thiago Serra, Srikumar Ramalingam, Shandian ZheICML 2022 · 60 citations
Related papers
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 440 citations
- SlimGPT: Layer-wise Structured Pruning for Large Language ModelsGui Ling, Ziyang Wang, Yuliang Yan, Qingwen LiuNeurIPS 2024 · 58 citations
- OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial OptimizationXiang Meng, Shibal Ibrahim, Kayhan Behdin, Hussein Hazimeh et al.ICML 2024 · 17 citations
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMsHang Guo, Luca Benini, Yawei LiICLR 2026 · 6 citations
- Hybrid Network Compression via Meta-LearningJianming Ye, Shiliang Zhang, Jingdong WangACM MM 2021 · 7 citations
