PackQViT: Faster Sub-8-bit Vision Transformers via Full and Packed Quantization on the Mobile
Peiyan Dong, Lei Lu, Chao Wu, Cheng Lyu, Geng Yuan, Hao Tang, Yanzhi Wang
Abstract
While Vision Transformers (ViTs) have undoubtedly made impressive strides in computer vision (CV), their intricate network structures necessitate substantial computation and memory resources. A decision-making process for CV tasks typically entails performing computations with low latency, which is a tricky problem for ViT models. Model quantization is a widely-used technique to optimize the hardware efficiency of deep neural networks. Full quantization under Sub-8-bit precision, in particular, is a promising solution to reduce inference latency significantly. Unfortunately, current commodity hardware, such as CPUs and GPUs, still struggles to efficiently execute these sub-8-bit quantized networks, as their SIMD instructions only support a granularity of 8 bits or wider. Also, there is a scarcity of literature that presents a full quantization paradigm for ViTs. In this paper, we propose an activation-aware fully sub-8-bit quantization-aware training (QAT) framework called PackQViT for efficient yet accurate ViT acceleration on mobile devices to facilitate real-time AI-powered decision-making. Specifically, in revisiting data activation within the ViT dataflow, two characteristics are relevant to quantization strategy and precision: the long-tailed distribution and systematic channel-wise outliers. In response, we employ either log2 quantization or clipping to address the long-tailed distribution and incorporate outlier-aware training for residual link quantization to regulate the various channel-wise outliers more consistently. Notably, due to the systematic fixed pattern, outlier-aware training approach can predict the channel indices and regularized scales of outliers in advance, thus avoiding the runtime data-adaptive selection during inference. Furthermore, we employ Int-2 n -Softmax, Int-LayerNorm, and Integer GELU to enable integer-only computation flow. Finally, we develop a SIMD-based 4-bit packed multiplier to achieve end-to-end ViT acceleration on mobile phones. Compared to prior studies on ViT quantization using 8-bit precision, PackQViT surpasses other works by an improved accuracy ranging from 0.4% to 17.9% for various widely used ViTs on ImageNet dataset; under 4-bit precision, PackQViT demonstrates 0.4%⇠2.8% higher accuracy. Compared to the baseline multiplier, our implementations on the Realme GT Android smartphone with Snapdragon 870 SoC CPU achieve 2.6 ⇥ ⇠3.7⇥ speedup under 8-bit scenario and 3.8 ⇥ ⇠5.9⇥ speedup under 4-bit which ensures practical real-time performance. Codes available at PackQViT. Value after GELU in FFN in 12-th block Attention Map with Quantization Level Distribution Value Frequency Weight Distribution of 12-th FC1 in DeiT-T Value Frequency
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cf04806-842f-442b-b461-a1706d5812faCited by top-tier papers9
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang et al.NeurIPS 2024 · 23 citations
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 10 citations
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model ServingJungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho et al.MICRO 2025 · 7 citations
- Ouromamba: a Data-Free Quantization Framework for Vision MambaAkshat Ramachandran, Mingyu Lee, Huan Xu, Souvik Kundu et al.ICCV 2025 · 3 citations
- PARO: Hardware-Software Co-design with Pattern-aware Reorder-based Attention Quantization in Video Generation ModelsXinhao Yang, Tianchen Zhao, Hongyi Wang, Wenheng Ma et al.DAC 2025 · 2 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
Related papers
- GradQ-ViT: Robust and Efficient Gradient Quantization for Vision TransformersDahun Choi, Hyun KimAAAI 2025 · 11 citations
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 172 citations
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie et al.HPCA 2023 · 117 citations
- UQ-ViT: Harmonizing Extreme Activations with Hardware-Friendly Uniform Quantization in Vision TransformersTao Jiang, Yucheng Jiang, Xiwen Yao, Gong Cheng et al.AAAI 2026
- QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceXinkuang Geng, Siting Liu, Leibo Liu, Jie Han et al.DAC 2024 · 5 citations
