MDP: Multidimensional Vision Model Pruning with Latency Constraint
Xinglong Sun, Barath Lakshmanan, Maying Shen, Shiyi Lan, Jingde Chen, José M. Álvarez
Abstract
Figure 1. MDP exhibits Pareto dominance with both CNNs and Transformers in tasks ranging from ImageNet classification to NuScenes 3D detection. Speedup are shown relative to the dense model. [Left] On ImageNet pruning ResNet50, we achieve a 28% speed increase alongside a +1.4 improvement in Top-1 compared with prior art [58]. [Middle] On ImageNet pruning DEIT-Base, compared with very recent Isomorphic Pruning (ECCV'24)[23], our method further accelerates the baseline by an additional 37% while yielding a +0.7 gain in Top-1. [Right] For 3D object detection, we obtain higher speed (×1.18) and mAP (0.451 vs. 0.449) compared to the dense baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73c96c7d-47b0-4be8-baf0-635c7b6b3faeCited by top-tier papers3
- NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge DevicesZiteng Wei, Qiang He, Bing Li, Feifei Chen et al.CVPR 2026 · 1 citation
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligenceZiteng Wei, Qiang He, Feifei Chen, Ranjie Duan et al.ICLR 2026
- CIGMA: Causal Information-Gain Mechanistic Attribution of Attention Heads in Vision TransformersMaisha Maliha, Dean F. HougenCVPR 2026
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
Related papers
- PDP: Parameter-free Differentiable Pruning is All You NeedMinsik Cho, Saurabh Adya, Devang NaikNeurIPS 2023 · 24 citations
- Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware PerspectiveZhenfeng Su, Kang Zhao, Han Bao, Tao Yuan et al.ICML 2026
- How Well Do Sparse ImageNet Models Transfer?Eugenia Iofinova, Alexandra Peste, Mark Kurtz, Dan AlistarhCVPR 2022 · 20 citations
- Effective Model Sparsification by Scheduled Grow-and-Prune MethodsXiaolong Ma, Minghai Qin, Fei Sun, Zejiang Hou et al.ICLR 2022 · 45 citations
- Manifold Regularized Dynamic Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Yiping Deng et al.CVPR 2021
