SAViT: Structure-Aware Vision Transformer Pruning via Collaborative Optimization
Chuanyang Zheng, Zheyang Li, Kai Zhang, Zhi Yang, Wenming Tan, Jun Xiao, Ye Ren, Shiliang Pu
Abstract
Vision Transformers (ViTs) yield impressive performance across various vision tasks. However, heavy computation and memory footprint make them inaccessible for edge devices. Previous works apply importance criteria determined independently by each individual component to prune ViTs. Considering that heterogeneous components in ViTs play distinct roles, these approaches lead to suboptimal performance. In this paper, we introduce joint importance, which integrates essential structural-aware interactions between components for the first time, to perform collaborative pruning. Based on the theoretical analysis, we construct a Taylor-based approximation to evaluate the joint importance. This guides pruning toward a more balanced reduction across all components. To further reduce the algorithm complexity, we incorporate the interactions into the optimization function under some mild assumptions. Moreover, the proposed method can be seamlessly applied to various tasks including object detection. Extensive experiments demonstrate the effectiveness of our method. Notably, the proposed approach outperforms the existing state-of-the-art approaches on ImageNet, increasing accuracy by 0.7% over the DeiT-Base baseline while saving 50% FLOPs. On COCO, we are the first to show that 70% FLOPs of Faster R-CNN with ViT backbone can be removed with only 0.3% mAP drop. The code is available at https://github.com/hikvision-research/SAViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0096d191-33df-48a7-b8f9-a0370675f71dCited by top-tier papers22
- Rep ViT: Revisiting Mobile CNN From ViT PerspectiveAo Wang, Hui Chen, Zijia Lin, Jungong Han et al.CVPR 2024 · 500 citations
- LGViT: Dynamic Early Exiting for Accelerating Vision TransformerGuanyu Xu, Jiawei Hao, Li Shen, Han Hu et al.ACM MM 2023 · 33 citations
- FDViT: Improve the Hierarchical Architecture of Vision TransformerYixing Xu, Chao Li, Dong Li, Xiao Sheng et al.ICCV 2023 · 21 citations
- The Need for Speed: Pruning Transformers with One RecipeSamir Khaki, Konstantinos N. PlataniotisICLR 2024 · 17 citations
- Data-independent Module-aware Pruning for Hierarchical Vision TransformersYang He, Joey Tianyi ZhouICLR 2024 · 11 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Synergistic Patch Pruning for Vision Transformer: Unifying Intra- & Inter-Layer Patch ImportanceYuyao Zhang, Lan Wei, Nikolaos M. FrerisICLR 2024 · 7 citations
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- ToaSt: Token Channel Selection and Structured Pruning for Efficient ViTHyunchan Moon, Cheonjun Park, Steven WaslanderICML 2026
- Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware PerspectiveZhenfeng Su, Kang Zhao, Han Bao, Tao Yuan et al.ICML 2026
