The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
Diyuan Wu, Ionut-Vlad Modoranu, Mher Safaryan, Denis Kuznedelev, Dan Alistarh
Abstract
The rising footprint of machine learning has led to a focus on imposing model sparsity as a means of reducing computational and memory costs. For deep neural networks (DNNs), the state-of-the-art accuracy-vs-sparsity is achieved by heuristics inspired by the classical Optimal Brain Surgeon (OBS) framework , which leverages loss curvature information to make better pruning decisions. Yet, these results still lack a solid theoretical understanding, and it is unclear whether they can be improved by leveraging connections to the wealth of work on sparse recovery algorithms. In this paper, we draw new connections between these two areas and present new sparse recovery algorithms inspired by the OBS framework that comes with theoretical guarantees under reasonable assumptions and have strong practical performance. Specifically, our work starts from the observation that we can leverage curvature information in OBS-like fashion upon the projection step of classic iterative sparse recovery algorithms such as IHT. We show for the first time that this leads both to improved convergence bounds under standard assumptions. Furthermore, we present extensions of this approach to the practical task of obtaining accurate sparse DNNs, and validate it experimentally at scale for Transformer-based models on vision and language tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b449beeb-fda2-4f81-b3dd-9ce175c1d90dCited by top-tier papers2
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-ShotKaiwen TUO, Huan WangICML 2026 · 6 citations
- WeightLoRA: Keep Only Necessary AdaptersAndrey Veprikov, Vladimir Solodkin, Alexander Zyl, Andrey V. Savchenko et al.ACL 2026 · 2 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 440 citations
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong et al.ICML 2024 · 306 citations
Related papers
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang et al.ICLR 2026 · 26 citations
- The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural NetworksXin Yu, Thiago Serra, Srikumar Ramalingam, Shandian ZheICML 2022 · 60 citations
- A Recovery Guarantee for Sparse Neural NetworksSara Fridovich-Keil, Mert PilanciICLR 2026 · 1 citation
- Optimal Eye Surgeon: Finding image priors through sparse generators at initializationAvrajit Ghosh, Xitong Zhang, Kenneth K. Sun, Qing Qu et al.ICML 2024 · 6 citations
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMsHang Guo, Luca Benini, Yawei LiICLR 2026 · 6 citations
