Understanding The Robustness in Vision Transformers
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, José M. Álvarez
Abstract
Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code is available at https://github.com/NVlabs/FAN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebfa086f-9ecd-4770-ad35-3e334cdc6af3Cited by top-tier papers77
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- Scaling & Shifting Your Features: A New Baseline for Efficient Model TuningDongze Lian, Daquan Zhou, Jiashi Feng, Xinchao WangNeurIPS 2022 · 415 citations
- Dual Aggregation Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong et al.ICCV 2023 · 345 citations
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 313 citations
- Scale-Aware Modulation Meet TransformerWeifeng Lin, Ziheng Wu, Jiayu Chen, Jun Huang et al.ICCV 2023 · 154 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Fully Attentional Networks with Self-emerging Token LabelingBingyin Zhao, Zhiding Yu, Shiyi Lan, Yutao Cheng et al.ICCV 2023 · 7 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Towards Robust Vision TransformerXiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li et al.CVPR 2022 · 185 citations
- Vision Transformers Are Robust LearnersSayak Paul, Pin-Yu ChenAAAI 2022 · 372 citations
- ReMoE: Region-Mixture Experts for Adversarially-Robust Vision TransformersQinghao Zhong, Bingzhi Chen, Yishu Liu, Minhua Lu et al.CVPR 2026
