Jumping through Local Minima: Quantization in the Loss Landscape of Vision Transformers
Natalia Frumkin, Dibakar Gope, Diana Marculescu
摘要
Quantization scale and bit-width are the most important parameters when considering how to quantize a neural network. Prior work focuses on optimizing quantization scales in a global manner through gradient methods (gradient descent & Hessian analysis). Yet, when applying perturbations to quantization scales, we observe a very jagged, highly non-smooth test loss landscape. In fact, small perturbations in quantization scale can greatly affect accuracy, yielding a 0.5 -0.8% accuracy boost in 4-bit quantized vision transformers (ViTs). In this regime, gradient methods break down, since they cannot reliably reach local minima. In our work, dubbed Evol-Q, we use evolutionary search to effectively traverse the non-smooth landscape. Additionally, we propose using an infoNCE loss, which not only helps combat overfitting on the small calibration dataset (1, 000 images) but also makes traversing such a highly non-smooth surface easier. Evol-Q improves the top-1 accuracy of a fully quantized ViT-Base by 10.30%, 0.78%, and 0.15% for 3-bit, 4-bit, and 8-bit weight quantization levels. Extensive experiments on a variety of CNN and ViT architectures further demonstrate its robustness in extreme quantization scenarios. Our code is available at https: // github . com/ enyac-group/ evol-q .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- PTQ4DiT: Post-training Quantization for Diffusion TransformersJunyi Wu, Haoxuan Wang, Yuzhang Shang, Mubarak Shah 等NeurIPS 2024 · 被引用 87 次
- Outlier-aware Slicing for Post-Training Quantization in Vision TransformerYuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling 等ICML 2024 · 被引用 17 次
- Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN InferenceAkshat Ramachandran, Zishen Wan, Geonhwa Jeong, John L. Gustafson 等DAC 2024 · 被引用 17 次
- ERQ: Error Reduction for Post-Training Quantization of Vision TransformersYunshan Zhong, Jiawei Hu, You Huang, Yuxin Zhang 等ICML 2024 · 被引用 14 次
- GradQ-ViT: Robust and Efficient Gradient Quantization for Vision TransformersDahun Choi, Hyun KimAAAI 2025 · 被引用 11 次
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceXinkuang Geng, Siting Liu, Leibo Liu, Jie Han 等DAC 2024 · 被引用 5 次
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao 等NeurIPS 2022 · 被引用 185 次
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 被引用 172 次
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 被引用 10 次
- AIQViT: Architecture-Informed Post-Training Quantization for Vision TransformersRunqing Jiang, Ye Zhang, Longguang Wang, Pengpeng Yu 等AAAI 2025 · 被引用 4 次
