Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning
Elias Frantar, Dan Alistarh
Abstract
We consider the problem of model compression for deep neural networks (DNNs) in the challenging one-shot/post-training setting, in which we are given an accurate trained model, and must compress it without any retraining, based only on a small amount of calibration input data. This problem has become popular in view of the emerging software and hardware support for executing models compressed via pruning and/or quantization with speedup, and well-performing solutions have been proposed independently for both compression approaches. In this paper, we introduce a new compression framework which covers both weight pruning and quantization in a unified setting, is time- and space-efficient, and considerably improves upon the practical performance of existing post-training methods. At the technical level, our approach is based on an exact and efficient realization of the classical Optimal Brain Surgeon (OBS) framework of [LeCun, Denker, and Solla, 1990] extended to also cover weight quantization at the scale of modern DNNs. From the practical perspective, our experimental results show that it can improve significantly upon the compression-accuracy trade-offs of existing post-training methods, and that it can enable the accurate compound application of both pruning and quantization in a post-training setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3b39fed-5ce0-4587-b584-127607af3e02Cited by top-tier papers125
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight CompressionTim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev et al.ICLR 2024 · 392 citations
- SliceGPT: Compress Large Language Models by Deleting Rows and ColumnsSaleh Ashkboos, Maximilian L. Croci, Marcelo Gennari Do Nascimento, Torsten Hoefler et al.ICLR 2024 · 346 citations
Builds on16
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
Related papers
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang et al.ICLR 2026 · 26 citations
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- CrAM: A Compression-Aware MinimizerAlexandra Peste, Adrian Vladu, Eldar Kurtic, Christoph H. Lampert et al.ICLR 2023 · 1 citation
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
