Boost Vision Transformer with GPU-Friendly Sparsity and Quantization
Chong Yu, Tao Chen, Zhongxue Gan, Jiayuan Fan
Abstract
The transformer extends its success from the language to the vision domain. Because of the stacked self-attention and cross-attention blocks, the acceleration deployment of vision transformer on GPU hardware is challenging and also rarely studied. This paper thoroughly designs a compression scheme to maximally utilize the GPU-friendly 2:4 finegrained structured sparsity and quantization. Specially, an original large model with dense weight parameters is first pruned into a sparse one by 2:4 structured pruning, which considers the GPU's acceleration of 2:4 structured sparse pattern with FP16 data type, then the floating-point sparse model is further quantized into a fixed-point one by sparsedistillation-aware quantization aware training, which considers GPU can provide an extra speedup of 2:4 sparse calculation with integer tensors. A mixed-strategy knowledge distillation is used during the pruning and quantization process. The proposed compression scheme is flexible to support supervised and unsupervised learning styles. Experiment results show GPUSQ-ViT scheme achieves state-ofthe-art compression by reducing vision transformer models 6.4-12.7× on model size and 30.3-62× on FLOPs with negligible accuracy degradation on ImageNet classification, COCO detection and ADE20K segmentation benchmarking tasks. Moreover, GPUSQ-ViT can boost actual deployment performance by 1.39-1.79× and 3.22-3.43× of latency and throughput on A100 GPU, and 1.57-1.69× and 2.11-2.51× improvement of latency and throughput on AGX Orin. Sparse M✕N✕K GEMM Dense M✕N✕K GEMM K A matrix (Dense) ☓ Accumulator (result) N Dense operation on Tensor Core M B matrix (Dense) C matrix (Dense) M K K/2 A matrix (Sparse) Non-zero data values 2-bits indices K/2 ☓ Accumulator (result) Sparse operation on Tensor Core Select B matrix (Dense) C matrix (Dense) N Choose matching K/2 elements out of K elements M M K
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim et al.ICLR 2026 · 5 citations
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision TransformersMinjun Kim, Jaeri Lee, Jongjin Kim, Jeongin Yun et al.AAAI 2026 · 1 citation
- Effective Interplay between Sparsity and Quantization: From Theory to PracticeSimla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin et al.ICLR 2025
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Deep Compression of Pre-trained Transformer ModelsNaigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen et al.NeurIPS 2022 · 38 citations
- CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision ModelsDenis Kuznedelev, Eldar Kurtic, Elias Frantar, Dan AlistarhNeurIPS 2023 · 24 citations
- SynGPU: Synergizing CUDA and Bit-Serial Tensor Cores for Vision Transformer Acceleration on GPUYuanzheng Yao, Chen Zhang, Chunyu Qi, Ruiyang Chen et al.DAC 2025 · 1 citation
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao et al.NeurIPS 2022 · 185 citations
- QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceXinkuang Geng, Siting Liu, Leibo Liu, Jie Han et al.DAC 2024 · 5 citations
