CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Models
Denis Kuznedelev, Eldar Kurtic, Elias Frantar, Dan Alistarh
Abstract
Driven by significant improvements in architectural design and training pipelines, computer vision has recently experienced dramatic progress in terms of accuracy on classic benchmarks such as ImageNet. These highly-accurate models are challenging to deploy, as they appear harder to compress using standard techniques such as pruning. We address this issue by introducing the Correlation Aware Pruner (CAP), a new unstructured pruning framework which significantly pushes the compressibility limits for state-of-the-art architectures. Our method is based on two technical advancements: a new theoretically-justified pruner, which can handle complex weight correlations accurately and efficiently during the pruning process itself, and an efficient finetuning procedure for post-compression recovery. We validate our approach via extensive experiments on several modern vision models such as Vision Transformers (ViT), modern CNNs, and ViT-CNN hybrids, showing for the first time that these can be pruned to high sparsity levels (e.g. %) with low impact on accuracy (% relative drop). Our approach is also compatible with structured pruning and quantization, and can lead to practical speedups of 1.5 to 2.4x without accuracy loss. To further showcase CAP's accuracy and scalability, we use it to show for the first time that extremely-accurate large vision models, trained via self-supervised techniques, can also be pruned to moderate sparsities, with negligible accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c4f8c81-1a7f-49f1-865a-91eee0e07d39Cited by top-tier papers7
- Hyperbolic Aware Minimization: Implicit Bias for SparsityTom Jacobs, Advait Gadhikar, Celia Rubio-Madrigal, Rebekka BurkholzICLR 2026 · 3 citations
- SPADE: Sparsity-Guided Debugging for Deep Neural NetworksArshia Soltani Moakhar, Eugenia Iofinova, Elias Frantar, Dan AlistarhICML 2024 · 2 citations
- SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial OptimizationTaisuke Yasuda, Kyriakos Axiotis, Gang Fu, Mohammad Hossein Bateni et al.NeurIPS 2024 · 1 citation
- Sign-In to the Lottery: Reparameterizing Sparse TrainingAdvait Gadhikar, Tom Jacobs, Chao Zhou, Rebekka BurkholzNeurIPS 2025
- Preserving Deep Representations in One-Shot Pruning: A Hessian-Free Second-Order Optimization FrameworkRyan Lucas, Rahul MazumderICLR 2025
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision TransformerMiao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin et al.AAAI 2023 · 23 citations
- Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware PerspectiveZhenfeng Su, Kang Zhao, Han Bao, Tao Yuan et al.ICML 2026
- Boost Vision Transformer with GPU-Friendly Sparsity and QuantizationChong Yu, Tao Chen, Zhongxue Gan, Jiayuan FanCVPR 2023
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
