Effective Interplay between Sparsity and Quantization: From Theory to Practice
Simla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin, Dongho Ha, Babak Falsafi, Martin Jaggi, Ming Liu, Yunho Oh, Suvinay Subramanian, Amir Yazdanbakhsh
Abstract
The increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints. Sparsity and quantization are two prominent compression methods that have been shown to reduce DNNs' computational and memory footprints significantly while preserving model accuracy. However, how these two methods interact when combined together remains a key question for developers, as many tacitly assume that they are orthogonal, meaning that their combined use does not introduce additional errors beyond those introduced by each method independently. In this paper, we provide the first mathematical proof that sparsity and quantization are non-orthogonal. We corroborate these results with experiments spanning a range of large language models, including the OPT and LLaMA model families (with 125M to 8B parameters), and vision models like ViT and ResNet. We show that the order in which we apply these methods matters because applying quantization before sparsity may disrupt the relative importance of tensor elements, which may inadvertently remove significant elements from a tensor. More importantly, we show that even if applied in the correct order, the compounded errors from sparsity and quantization can significantly harm accuracy. Our findings extend to the efficient deployment of large models in resource-constrained compute platforms to reduce serving cost, offering insights into best practices for applying these compression methods to maximize hardware resource efficiency without compromising accuracy. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41450659-df7d-4b17-b9b7-5159eb7478d0Cited by top-tier papers7
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou et al.NeurIPS 2024 · 47 citations
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 12 citations
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMsHang Guo, Luca Benini, Yawei LiICLR 2026 · 6 citations
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim et al.ICLR 2026 · 5 citations
- Unified Scaling Laws for Compressed RepresentationsAndrei Panferov, Alexandra Volkova, Ionut-Vlad Modoranu, Vage Egiazarian et al.NeurIPS 2025 · 5 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
Related papers
- Pruning vs Quantization: Which is Better?Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi et al.NeurIPS 2023 · 152 citations
- Compressing Large Language Models by Joint Sparsification and QuantizationJinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu et al.ICML 2024 · 33 citations
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
- REx: Data-Free Residual Quantization Error ExpansionEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2023 · 11 citations
- Permute, Quantize, and Fine-Tune: Efficient Compression of Neural NetworksJulieta Martinez, Jashan Shewakramani, Ting-Wei Liu, Ioan Andrei Barsan et al.CVPR 2021
