SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
Minjun Kim, Jongjin Kim, U Kang
Abstract
How can we accurately quantize a pre-trained model without any data? Quantization algorithms are widely used for deploying neural networks on resource-constrained edge devices. Zero-shot Quantization (ZSQ) addresses the crucial and practical scenario where training data are inaccessible for privacy or security reasons. However, three significant challenges hinder the performance of existing ZSQ methods: 1) noise in the synthetic dataset, 2) predictions based on off-target patterns, and the 3) misguidance by erroneous hard labels. In this paper, we propose SYNQ (Synthesis-aware Fine-tuning for Zero-shot Quantization), a carefully designed ZSQ framework to overcome the limitations of existing methods. SYNQ minimizes the noise from the generated samples by exploiting a low-pass filter. Then, SYNQ trains the quantized model to improve accuracy by aligning its class activation map with the pre-trained model. Furthermore, SYNQ mitigates misguidance from the pre-trained model's error by leveraging only soft labels for difficult samples. Extensive experiments show that SYNQ provides the state-of-the-art accuracy, over existing ZSQ methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 909fdc18-07e0-453c-9be2-27d0c0621a59Cited by top-tier papers5
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim et al.ICLR 2026 · 5 citations
- Ouromamba: a Data-Free Quantization Framework for Vision MambaAkshat Ramachandran, Mingyu Lee, Huan Xu, Souvik Kundu et al.ICCV 2025 · 3 citations
- Semantic Alignment and Reinforcement for Data-Free Quantization of Vision TransformersYunshan Zhong, Yuyao Zhou, Yuxin Zhang, Wanchen Sui et al.ICCV 2025 · 2 citations
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision TransformersMinjun Kim, Jaeri Lee, Jongjin Kim, Jeongin Yun et al.AAAI 2026 · 1 citation
- Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision TransformersBiao Qian, Yang Wang, Yong Wu, Jungong HanICML 2026 · 1 citation
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
Related papers
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
- Hard Sample Matters a Lot in Zero-Shot QuantizationHuantong Li, Xiangmiao Wu, Fanbing Lv, Daihai Liao et al.CVPR 2023
- Task-Specific Zero-Shot Quantization-Aware Training for Object DetectionChanghao Li, Xinrui Chen, Ji Wang, Kang Zhao et al.ICCV 2025 · 2 citations
- Genie: Show Me the Data for QuantizationYongkweon Jeon, Chungman Lee, Ho-Young KimCVPR 2023
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu et al.NeurIPS 2023 · 24 citations
