Pruning vs Quantization: Which is Better?
Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi, Tijmen Blankevoort
摘要
Neural network pruning and quantization techniques are almost as old as neural networks themselves. However, to date only ad-hoc comparisons between the two have been published. In this paper, we set out to answer the question on which is better: neural network quantization or pruning? By answering this question, we hope to inform design decisions made on neural network hardware going forward. We provide an extensive comparison between the two techniques for compressing deep neural networks. First, we give an analytical comparison of expected quantization and pruning error for general data distributions. Then, we provide lower bounds for the per-layer pruning and quantization error in trained networks, and compare these to empirical error after optimization. Finally, we provide an extensive experimental comparison for training 9 large-scale models on 4 tasks. Our results show that in most cases quantization outperforms pruning. Only in some scenarios with very high compression ratio, pruning might be beneficial from an accuracy standpoint. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 被引用 21 次
- Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMsJordan Dotzel, Yuzong Chen, Bahaa Kotb, Sushma Prasad 等ICML 2024 · 被引用 20 次
- ERQ: Error Reduction for Post-Training Quantization of Vision TransformersYunshan Zhong, Jiawei Hu, You Huang, Yuxin Zhang 等ICML 2024 · 被引用 14 次
- MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling QuantizationAkshat Ramachandran, Souvik Kundu, Tushar KrishnaISCA 2025 · 被引用 10 次
- Large Language Model Compression with Global Rank and Sparsity OptimizationChanghai Zhou, Qian Qiao, Yuhua Zhou, Yuxin Wu 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- Effective Interplay between Sparsity and Quantization: From Theory to PracticeSimla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin 等ICLR 2025
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 被引用 440 次
- A Probabilistic Approach to Neural Network PruningXin Qian, Diego KlabjanICML 2021 · 被引用 24 次
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian 等ICCV 2023 · 被引用 31 次
