Augmented Shortcuts for Vision Transformers
Yehui Tang, Kai Han, Chang Xu, An Xiao, Yiping Deng, Chao Xu, Yunhe Wang
Abstract
Transformer models have achieved great progress on computer vision tasks recently. The rapid development of vision transformers is mainly contributed by their high representation ability for extracting informative features from input images. However, the mainstream transformer models are designed with deep architectures, and the feature diversity will be continuously reduced as the depth increases, i.e., feature collapse. In this paper, we theoretically analyze the feature collapse phenomenon and study the relationship between shortcuts and feature diversity in these transformer models. Then, we present an augmented shortcut scheme, which inserts additional paths with learnable parameters in parallel on the original shortcuts. To save the computational costs, we further explore an efficient approach that uses the block-circulant projection to implement augmented shortcuts. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed method, which brings about 1% accuracy increase of the state-of-the-art visual transformers without obviously increasing their parameters and FLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc837169-4d31-4e9c-87b9-9eb3ebcf99cbCited by top-tier papers11
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- GhostNetV2: Enhance Cheap Operation with Long-Range AttentionYehui Tang, Kai Han, Jianyuan Guo, Chang Xu et al.NeurIPS 2022 · 634 citations
- Multi-Scale High-Resolution Vision Transformer for Semantic SegmentationJiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye et al.CVPR 2022 · 236 citations
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 212 citations
- Patch Slimming for Efficient Vision TransformersYehui Tang, Kai Han, Yunhe Wang, Chang Xu et al.CVPR 2022 · 173 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- Mean-Shift Feature TransformerTakumi KobayashiCVPR 2024
- Frequency-Aware Token Reduction for Efficient Vision TransformerDongJae Lee, Jiwan Hur, Jaehyun Choi, Jaemyung Yu et al.NeurIPS 2025 · 4 citations
- Visual Transformer with Differentiable Channel Selection: An Information Bottleneck Inspired ApproachYancheng Wang, Ping Li, Yingzhen YangICML 2024 · 2 citations
- Width & Depth Pruning for Vision TransformersFang Yu, Kun Huang, Meng Wang, Yuan Cheng et al.AAAI 2022 · 159 citations
- Adder Attention for Vision TransformerHan Shu, Jiahao Wang, Hanting Chen, Lin Li et al.NeurIPS 2021 · 23 citations
