Hard Sample Matters a Lot in Zero-Shot Quantization
Huantong Li, Xiangmiao Wu, Fanbing Lv, Daihai Liao, Thomas H. Li, Yonggang Zhang, Bo Han, Mingkui Tan
Abstract
Zero-shot quantization (ZSQ) is promising for compressing and accelerating deep neural networks when the data for training full-precision models are inaccessible. In ZSQ, network quantization is performed using synthetic samples, thus, the performance of quantized models depends heavily on the quality of synthetic samples. Nonetheless, we find that the synthetic samples constructed in existing ZSQ methods can be easily fitted by models. Accordingly, quantized models obtained by these methods suffer from significant performance degradation on hard samples. To address this issue, we propose HArd sample Synthesizing and Training (HAST). Specifically, HAST pays more attention to hard samples when synthesizing samples and makes synthetic samples hard to fit when training quantized models. HAST aligns features extracted by full-precision and quantized models to ensure the similarity between features extracted by these two models. Extensive experiments show that HAST significantly outperforms existing ZSQ methods, achieving performance comparable to models that are quantized with real data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8f144c2-7cd9-486c-8038-877da4916342Cited by top-tier papers9
- Enhancing One-Shot Federated Learning Through Data and Ensemble Co-BoostingRong Dai, Yonggang Zhang, Ang Li, Tongliang Liu et al.ICLR 2024 · 40 citations
- SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin RegularizationChen Tang, Kai Ouyang, Zenghao Chai, Yunpeng Bai et al.ACM MM 2023 · 11 citations
- Sharpness-Aware Data Generation for Zero-shot QuantizationHoang Anh Dung, Cuong Pham, Trung Le, Jianfei Cai et al.ICML 2024 · 8 citations
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 7 citations
- Semantic Alignment and Reinforcement for Data-Free Quantization of Vision TransformersYunshan Zhong, Yuyao Zhou, Yuxin Zhang, Wanchen Sui et al.ICCV 2025 · 2 citations
Builds on23
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 315 citations
Related papers
- SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuningMinjun Kim, Jongjin Kim, U KangICLR 2025
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
- The Knowledge Within: Methods for Data-Free Model CompressionMatan Haroush, Itay Hubara, Elad Hoffer, Daniel SoudryCVPR 2020
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu et al.NeurIPS 2023 · 24 citations
- Genie: Show Me the Data for QuantizationYongkweon Jeon, Chungman Lee, Ho-Young KimCVPR 2023
