Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice
Peihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang Wang
Abstract
Vision Transformer (ViT) has recently demonstrated promise in computer vision problems. However, unlike Convolutional Neural Networks (CNN), it is known that the performance of ViT saturates quickly with depth increasing, due to the observed attention collapse or patch uniformity. Despite a couple of empirical solutions, a rigorous framework studying on this scalability issue remains elusive. In this paper, we first establish a rigorous theory framework to analyze ViT features from the Fourier spectrum domain. We show that the self-attention mechanism inherently amounts to a low-pass filter, which indicates when ViT scales up its depth, excessive low-pass filtering will cause feature maps to only preserve their Direct-Current (DC) component. We then propose two straightforward yet effective techniques to mitigate the undesirable low-pass limitation. The first technique, termed AttnScale, decomposes a self-attention block into low-pass and high-pass components, then rescales and combines these two filters to produce an all-pass self-attention matrix. The second technique, termed FeatScale, re-weights feature maps on separate frequency bands to amplify the high-frequency signals. Both techniques are efficient and hyperparameter-free, while effectively overcoming relevant ViT training artifacts such as attention collapse and patch uniformity. By seamlessly plugging in our techniques to multiple ViT variants, we demonstrate that they consistently help ViTs benefit from deeper architectures, bringing up to 1.1% performance gains "for free" (e.g., with little parameter overhead). We publicly release our codes and pre-trained models at https://github.com/VITA-Group/ViT-Anti-Oversmoothing .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6696ad29-1e36-47c9-aa79-e7a2ee199a16Cited by top-tier papers69
- M³ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-designHanxue Liang, Zhiwen Fan, Rishov Sarkar, Ziyu Jiang et al.NeurIPS 2022 · 152 citations
- MogaNet: Multi-order Gated Aggregation NetworkSiyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan et al.ICLR 2024 · 151 citations
- An Attentive Inductive Bias for Sequential Recommendation beyond the Self-AttentionYehjin Shin, Jeongwhan Choi, Hyowon Wi, Noseong ParkAAAI 2024 · 130 citations
- Vision HGNN: An Image is More than a Graph of NodesYan Han, Peihao Wang, Souvik Kundu, Ying Ding et al.ICCV 2023 · 86 citations
- Fredformer: Frequency Debiased Transformer for Time Series ForecastingXihao Piao, Zheng Chen, Taichi Murayama, Yasuko Matsubara et al.KDD 2024 · 73 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 13 citations
- The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyTianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah et al.CVPR 2022 · 36 citations
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
- VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsZhenyu Wang, Hao Luo, Pichao Wang, Feng Ding et al.NeurIPS 2022 · 57 citations
- You Only Need Less Attention at Each Stage in Vision TransformersShuoxi Zhang, Hanpeng Liu, Stephen Lin, Kun HeCVPR 2024 · 19 citations
