Robust Quantization: One Model to Rule Them All
Moran Shkolnik, Brian Chmiel, Ron Banner, Gil Shomron, Yury Nahshan, Alex M. Bronstein, Uri C. Weiser
Abstract
Neural network quantization methods often involve simulating the quantization process during training, making the trained model highly dependent on the target bit-width and precise way quantization is performed. Robust quantization offers an alternative approach with improved tolerance to different classes of data-types and quantization policies. It opens up new exciting applications where the quantization process is not static and can vary to meet different circumstances and implementations. To address this issue, we propose a method that provides intrinsic robustness to the model against a broad range of quantization processes. Our method is motivated by theoretical arguments and enables us to store a single generic model capable of operating at various bit-widths and quantization policies. We validate our method's effectiveness on different ImageNet models. A reference implementation accompanies the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abcb7c2e-613c-49e7-86ce-447bf52a1b24Cited by top-tier papers32
- Outlier Suppression: Pushing the Limit of Low-bit Transformer Language ModelsXiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong et al.NeurIPS 2022 · 238 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
- Degree-Quant: Quantization-Aware Training for Graph Neural NetworksShyam Anil Tailor, Javier Fernández-Marqués, Nicholas Donald LaneICLR 2021 · 180 citations
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchMingzhu Shen, Feng Liang, Ruihao Gong, Yuhang Li et al.ICCV 2021 · 50 citations
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu et al.ICML 2022 · 49 citations
Builds on2
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
Related papers
- Gradient Regularization for Quantization RobustnessMilad Alizadeh, Arash Behboodi, Mart van Baalen, Christos Louizos et al.ICLR 2020 · 8 citations
- Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference ApproachChen Tang, Haoyu Zhai, Kai Ouyang, Zhi Wang et al.ACM MM 2022 · 15 citations
- EQ-Net: Elastic Quantization Neural NetworksKe Xu, Lei Han, Ye Tian, Shangshang Yang et al.ICCV 2023 · 21 citations
- DSConv: Efficient Convolution OperatorMarcelo Gennari Do Nascimento, Victor Prisacariu, Roger FawcettICCV 2019 · 107 citations
- On Quantizing Neural Representation for Variable-Rate Video CodingJunqi Shi, Zhujia Chen, Hanfei Li, Qi Zhao et al.ICLR 2025
