VTC-LFC: Vision Transformer Compression with Low-Frequency Components
Zhenyu Wang, Hao Luo, Pichao Wang, Feng Ding, Fan Wang, Hao Li
Abstract
Although Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without finetuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ∼ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K. Code will be available at https://github.com/Daner-Wang/VTC-LFC.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2439c43c-700f-4a08-a3e7-231052fb6d73Cited by top-tier papers11
- DiffRate : Differentiable Compression Rate for Efficient Vision TransformersMengzhao Chen, Wenqi Shao, Peng Xu, Mingbao Lin et al.ICCV 2023 · 87 citations
- Good Helper Is around You: Attention-Driven Masked Image ModelingZhengqi Liu, Jie Gui, Hao LuoAAAI 2023 · 36 citations
- Sparse Imagination for Efficient Visual World Model PlanningJunha Chun, Youngjoon Jeong, Taesup KimICLR 2026 · 8 citations
- Synergistic Patch Pruning for Vision Transformer: Unifying Intra- & Inter-Layer Patch ImportanceYuyao Zhang, Lan Wei, Nikolaos M. FrerisICLR 2024 · 7 citations
- Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer CompressionHancheng Ye, Chong Yu, Peng Ye, Renqiu Xia et al.CVPR 2024 · 7 citations
Builds on42
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- TF-ATM: Training-Free Adaptive Token MergingXin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen et al.ACM MM 2025
- Frequency-Aware Token Reduction for Efficient Vision TransformerDongJae Lee, Jiwan Hur, Jaehyun Choi, Jaemyung Yu et al.NeurIPS 2025 · 4 citations
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan et al.ICLR 2022 · 118 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Dense Vision Transformer Compression with Few SamplesHanxiao Zhang, Yifan Zhou, Guo-Hua WangCVPR 2024 · 5 citations
