V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision Transformer
Guangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang, Renda Han, Chuangxin Zhao, Zeyu Zhang, Runhao Liu
Abstract
Vision Transformer (ViT) has become one of the cornerstones of the computer vision field, demonstrating exceptional performance. However, its inherent high computational complexity and inference latency still pose significant obstacles for deployment in resource-constrained environments. Token pruning, by removing less informative tokens, offers an effective strategy to reduce computational overhead. However, existing pruning methods largely rely on static or local token importance scores. This myopic approach fundamentally overlooks the sequential dependency of pruning decisions and fails to capture the interaction effects between pruning decisions across layers, often neglecting the global interactions between mask variables. To address this limitation, we propose V-Pruner, a fast and globally-informed token pruning framework for Vision Transformer. V-Pruner first leverages Fisher information to perform an initial assessment of token importance, providing a principled initial prior for pruning decisions. Building on this, V-Pruner introduces a Reinforcement Learning (RL) Proximal Policy Optimization (PPO) algorithm, refining token pruning into a global sequential decision process. The algorithm combines a composite reward signal that incorporates both model performance and computational cost to guide policy exploration, effectively evaluating the long-term impact of different pruning decision combinations on global model performance. Extensive experiments on ViT-L, DeiT-B, DeiT-S, and DeiT-T demonstrate that V-Pruner achieves a better balance between accuracy, GFLOPs, inference speed, and training time, surpassing existing mainstream ViT pruning algorithms in overall performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ad32d56-5b54-41b2-8f0c-74288f22d64fBuilds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Learned Token Pruning for TransformersSehoon Kim, Sheng Shen, David Thorsley, Amir Gholami et al.KDD 2022 · 97 citations
Related papers
- HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersPeiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie et al.HPCA 2023 · 117 citations
- Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision TransformerYifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng et al.AAAI 2022 · 288 citations
- TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsHengyi Wang, Weiying Xie, Hui Jiang, Yaotao Wei et al.AAAI 2026
- Multi-Criteria Token Fusion with One-Step-Ahead Attention for Efficient Vision TransformersSanghyeok Lee, Joonmyung Choi, Hyunwoo J. KimCVPR 2024
- Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersSifan Long, Zhen Zhao, Jimin Pi, Shengsheng Wang et al.CVPR 2023
