Synergistic Patch Pruning for Vision Transformer: Unifying Intra- & Inter-Layer Patch Importance
Yuyao Zhang, Lan Wei, Nikolaos M. Freris
Abstract
The Vision Transformer (ViT) has emerged as a powerful architecture for various computer vision tasks. Nonetheless, this comes with substantially heavier computational costs than Convolutional Neural Networks (CNNs). The attention mechanism in ViTs, which integrates information from different image patches to the class token ([CLS]), renders traditional structured pruning methods used in CNNs unsuitable. To overcome this issue, we propose SynergisTic pAtch pRuning (STAR) that unifies intra-layer and inter-layer patch importance scoring. Specifically, our approach combines a) online evaluation of intra-layer importance for the [CLS] and b) offline evaluation of the inter-layer importance of each patch. The two importance scores are fused by minimizing a weighted average of Kullback-Leibler (KL) Divergences and patches are successively pruned at each layer by maintaining only the top-k most important ones. Unlike prior art that relies on manual selection of the pruning rates at each layer, we propose an automated method for selecting them based on offline-derived metrics. We also propose a variant that uses these rates as weighted percentile parameters (for the layer-wise normalized scores), thus leading to an alternate adaptive rate selection technique that is input-based. Extensive experiments demonstrate the significant acceleration of the inference with minimal performance degradation. For instance, on the ImageNet dataset, the pruned DeiT-Small reaches a throughput of 4,256 images/s, which is over 66% higher than the much smaller (unpruned) DeiT-Tiny model, while having a substantially higher accuracy (+6.8% Top-1 and +3.1% Top-5).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3992e025-1800-4644-841a-79527fbb157aCited by top-tier papers2
- Sparse Imagination for Efficient Visual World Model PlanningJunha Chun, Youngjoon Jeong, Taesup KimICLR 2026 · 8 citations
- Energy Landscape-Aware Vision Transformers: Layerwise Dynamics and Adaptive Task-Specific Training via Hopfield StatesRunze Xia, Richard JiangNeurIPS 2025 · 1 citation
Builds on23
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie et al.ICCV 2021 · 723 citations
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang et al.NeurIPS 2021 · 528 citations
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 335 citations
Related papers
- SAViT: Structure-Aware Vision Transformer Pruning via Collaborative OptimizationChuanyang Zheng, Zheyang Li, Kai Zhang, Zhi Yang et al.NeurIPS 2022 · 98 citations
- Global Vision Transformer Pruning with Hessian-Aware SaliencyHuanrui Yang, Hongxu Yin, Maying Shen, Pavlo Molchanov et al.CVPR 2023
- V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision TransformerGuangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang et al.AAAI 2026 · 1 citation
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision TransformerMiao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin et al.AAAI 2023 · 23 citations
