Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
Hancheng Ye, Chong Yu, Peng Ye, Renqiu Xia, Yansong Tang, Jiwen Lu, Tao Chen, Bo Zhang
Abstract
Recent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme, where the importance score of each model unit is first evaluated or preset in each submodule, followed by the sparsity score evaluation according to the target sparsity constraint. Such a separate evaluation process induces the gap between importance and sparsity score distributions, thus causing high search costs for VTC. In this work, for the first time, we investigate how to integrate the evaluations of importance and sparsity scores into a single stage, searching the optimal subnets in an efficient manner. Specifically, we present OFB, a cost-efficient approach that simultaneously evaluates both importance and sparsity scores, termed Once for Both (OFB), for VTC. First, a bi-mask scheme is developed by entangling the importance score and the differentiable sparsity score to jointly determine the pruning potential (prunability) of each unit. Such a bi-mask search strategy is further used together with a proposed adaptive one-hot loss to realize the progressiveand-efficient search for the most important subnet. Finally, Progressive Masked Image Modeling (PMIM) is proposed to regularize the feature space to be more representative during the search process, which may be degraded by the dimension reduction. Extensive experiments demonstrate that OFB can achieve superior compression performance over state-ofthe-art searching-based and pruning-based methods under various Vision Transformer architectures, meanwhile promoting search efficiency significantly, e.g., costing one GPU search day for the compression of DeiT-S on ImageNet-1K. * The importance evaluation aims at learning each unit's contribution to the prediction performance, while the sparsity evaluation aims at learning each unit's pruning choice. In general, the importance and sparsity score distributions are correlated in the search process, as shown in Fig. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4dfd4ef-b434-4420-baf7-ec55e84eabc8Cited by top-tier papers5
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent SystemsHancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang et al.NeurIPS 2025 · 42 citations
- Training-Free Adaptive Diffusion with Bounded Difference Approximation StrategyHancheng Ye, Jiakang Yuan, Renqiu Xia, Xiangchao Yan et al.NeurIPS 2024 · 21 citations
- Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own SignalsQinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin et al.NeurIPS 2025 · 16 citations
- V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision TransformerGuangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang et al.AAAI 2026 · 1 citation
- Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and CompressionNazia Tasnim, Shrimai Prabhumoye, Bryan A. PlummerCVPR 2026
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Collaborative Multi-Mode Pruning for Vision-Language ModelsZimeng Wu, Yunhong Wang, Donghao Wang, Jiaxin ChenCVPR 2026 · 2 citations
- Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization SpaceArnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu et al.CVPR 2022 · 64 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Synergistic Patch Pruning for Vision Transformer: Unifying Intra- & Inter-Layer Patch ImportanceYuyao Zhang, Lan Wei, Nikolaos M. FrerisICLR 2024 · 7 citations
- SAViT: Structure-Aware Vision Transformer Pruning via Collaborative OptimizationChuanyang Zheng, Zheyang Li, Kai Zhang, Zhi Yang et al.NeurIPS 2022 · 98 citations
