It's All In the Teacher: Zero-Shot Quantization Brought Closer to the Teacher
Kanghyun Choi, Hyeyoon Lee, Deokki Hong, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee
摘要
Model quantization is considered as a promising method to greatly reduce the resource requirements of deep neural networks. To deal with the performance drop induced by quantization errors, a popular method is to use training data to fine-tune quantized networks. In real-world environments, however, such a method is frequently infeasible because training data is unavailable due to security, privacy, or confidentiality concerns. Zero-shot quantization addresses such problems, usually by taking information from the weights of a full-precision teacher network to compensate the performance drop of the quantized networks. In this paper, we first analyze the loss surface of state-of-the-art zero-shot quantization techniques and provide several findings. In contrast to usual knowledge distillation problems, zero-shot quantization often suffers from 1) the difficulty of optimizing multiple loss terms together, and 2) the poor generalization capability due to the use of synthetic samples. Furthermore, we observe that many weights fail to cross the rounding threshold during training the quantized networks even when it is necessary to do so for better performance. Based on the observations, we propose AIT, a simple yet powerful technique for zero-shot quantization, which addresses the aforementioned two problems in the following way: AIT i) uses a KL distance loss only without a cross-entropy loss, and ii) manipulates gradients to guarantee that a certain portion of weights are properly updated after crossing the rounding thresholds. Experiments show that AIT outperforms the performance of many existing methods by a great margin, taking over the overall state-of-the-art position in the field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu 等NeurIPS 2023 · 被引用 24 次
- Rethinking Data-Free Quantization as a Zero-Sum GameBiao Qian, Yang Wang, Richang Hong, Meng WangAAAI 2023 · 被引用 24 次
- REx: Data-Free Residual Quantization Error ExpansionEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2023 · 被引用 11 次
- MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention SimilarityKanghyun Choi, Hyeyoon Lee, Dain Kwon, Sunjong Park 等AAAI 2025 · 被引用 9 次
- Robustness-Guided Image Synthesis for Data-Free QuantizationJianhong Bai, Yuchen Yang, Huanpeng Chu, Hualiang Wang 等AAAI 2024 · 被引用 7 次
它引用的顶会 Paper11
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- Qimera: Data-free Quantization with Synthetic Boundary Supporting SamplesKanghyun Choi, Deokki Hong, Noseong Park, Youngsok Kim 等NeurIPS 2021 · 被引用 87 次
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg 等ICML 2021 · 被引用 78 次
相关 Paper
- Genie: Show Me the Data for QuantizationYongkweon Jeon, Chungman Lee, Ho-Young KimCVPR 2023
- Sharpness-Aware Data Generation for Zero-shot QuantizationHoang Anh Dung, Cuong Pham, Trung Le, Jianfei Cai 等ICML 2024 · 被引用 8 次
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
- Task-Specific Zero-Shot Quantization-Aware Training for Object DetectionChanghao Li, Xinrui Chen, Ji Wang, Kang Zhao 等ICCV 2025 · 被引用 2 次
- ZeroQ: A Novel Zero Shot Quantization FrameworkYaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami 等CVPR 2020
