Dense Vision Transformer Compression with Few Samples
Hanxiao Zhang, Yifan Zhou, Guo-Hua Wang
Abstract
Few-shot model compression aims to compress a large model into a more compact one with only a tiny training set (even without labels). Block-level pruning has recently emerged as a leading technique in achieving high accuracy and low latency in few-shot CNN compression. But, few-shot compression for Vision Transformers (ViT) remains largely unexplored, which presents a new challenge. In particular, the issue of sparse compression exists in traditional CNN few-shot methods, which can only produce very few compressed models of different model sizes. This paper proposes a novel framework for few-shot ViT compression named DC-ViT. Instead of dropping the entire block, DC-ViT selectively eliminates the attention module while retaining and reusing portions of the MLP module. DC-ViT enables dense compression, which outputs numerous compressed models that densely populate the range of model complexity. DC-ViT outperforms state-of-the-art few-shot compression methods by a significant margin of 10 percentage points, along with lower latency in the compression of ViT and its variants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0dc24c2-6f77-4b0b-bc80-c3879333ffdfCited by top-tier papers7
- NuWa: Deriving Lightweight Class-Specific Vision Transformers for Edge DevicesZiteng Wei, Qiang He, Bing Li, Feifei Chen et al.CVPR 2026 · 1 citation
- Stratified Knowledge-Density Super-Network for Scalable Vision TransformersLonghua Li, Lei Qi, Xin GengAAAI 2026 · 1 citation
- Memory-Efficient Generative Models via Product QuantizationJie Shao, Hanxiao Zhang, Hao Yu, Jianxin WuICCV 2025 · 1 citation
- Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model ApproachXu Zhang, Kaidi Xu, Ziqing Hu, Ren WangICML 2025
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligenceZiteng Wei, Qiang He, Feifei Chen, Ranjie Duan et al.ICLR 2026
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
Related papers
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- MiniViT: Compressing Vision Transformers with Weight MultiplexingJinnian Zhang, Houwen Peng, Kan Wu, Mengchen Liu et al.CVPR 2022 · 115 citations
- GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision TransformerMiao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin et al.AAAI 2023 · 23 citations
- VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsZhenyu Wang, Hao Luo, Pichao Wang, Feng Ding et al.NeurIPS 2022 · 57 citations
