SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
Minjun Kim, Jongjin Kim, U Kang
摘要
How can we accurately quantize a pre-trained model without any data? Quantization algorithms are widely used for deploying neural networks on resource-constrained edge devices. Zero-shot Quantization (ZSQ) addresses the crucial and practical scenario where training data are inaccessible for privacy or security reasons. However, three significant challenges hinder the performance of existing ZSQ methods: 1) noise in the synthetic dataset, 2) predictions based on off-target patterns, and the 3) misguidance by erroneous hard labels. In this paper, we propose SYNQ (Synthesis-aware Fine-tuning for Zero-shot Quantization), a carefully designed ZSQ framework to overcome the limitations of existing methods. SYNQ minimizes the noise from the generated samples by exploiting a low-pass filter. Then, SYNQ trains the quantized model to improve accuracy by aligning its class activation map with the pre-trained model. Furthermore, SYNQ mitigates misguidance from the pre-trained model's error by leveraging only soft labels for difficult samples. Extensive experiments show that SYNQ provides the state-of-the-art accuracy, over existing ZSQ methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim 等ICLR 2026 · 被引用 5 次
- Ouromamba: a Data-Free Quantization Framework for Vision MambaAkshat Ramachandran, Mingyu Lee, Huan Xu, Souvik Kundu 等ICCV 2025 · 被引用 3 次
- Semantic Alignment and Reinforcement for Data-Free Quantization of Vision TransformersYunshan Zhong, Yuyao Zhou, Yuxin Zhang, Wanchen Sui 等ICCV 2025 · 被引用 2 次
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision TransformersMinjun Kim, Jaeri Lee, Jongjin Kim, Jeongin Yun 等AAAI 2026 · 被引用 1 次
- Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision TransformersBiao Qian, Yang Wang, Yong Wu, Jungong HanICML 2026 · 被引用 1 次
它引用的顶会 Paper34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
- Hard Sample Matters a Lot in Zero-Shot QuantizationHuantong Li, Xiangmiao Wu, Fanbing Lv, Daihai Liao 等CVPR 2023
- Task-Specific Zero-Shot Quantization-Aware Training for Object DetectionChanghao Li, Xinrui Chen, Ji Wang, Kang Zhao 等ICCV 2025 · 被引用 2 次
- Genie: Show Me the Data for QuantizationYongkweon Jeon, Chungman Lee, Ho-Young KimCVPR 2023
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu 等NeurIPS 2023 · 被引用 24 次
