Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Training
Seongmin Kim, Byung Cheol Song
Abstract
The Vision Transformer (ViT) has surpassed Convolutional Neural Networks (CNNs) in performance, becoming the de facto architecture in modern computer vision. However, despite its superior representational capacity, research on the adversarial robustness of ViTs remains limited, with most studies still biased toward CNN-based models. This work aims to address this architectural bias and conduct an indepth analysis of the interaction between ViTs and adversarial training (AT). We first show that ViTs can identify semantic components of objects through their class attention maps, indicating that adversarially trained ViTs inherently encode strong semantic priors. Next, using the proposed Gradient Path Masking (GPM) analysis, we examine the internal information flow of ViTs and verify that the residual path serves as a major bottleneck that provides advantageous information to adversaries. Furthermore, our inter-patch relation analysis reveals that adversarially trained ViTs tend to rely more on global than local relationships in early layers-a novel observation suggesting a potential incompatibility between ViTs and hybrid architectures that inject CNN-style inductive biases. Building upon these findings, we design a simple yet effective two-stage AT scheme to mitigate this structural incompatibility, achieving simultaneous improvements in robustness and generalization across various ViT variants and training methods. The proposed method is compatible with a wide range of AT frameworks and models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b047745f-8e90-4501-b214-685e0fc512e1Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- When Adversarial Training Meets Vision Transformers: Recipes from Training to ArchitectureYichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo et al.NeurIPS 2022 · 109 citations
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu et al.AAAI 2022 · 156 citations
- ReMoE: Region-Mixture Experts for Adversarially-Robust Vision TransformersQinghao Zhong, Bingzhi Chen, Yishu Liu, Minhua Lu et al.CVPR 2026
- Towards Understanding and Improving Adversarial Robustness of Vision TransformersSamyak Jain, Tanima DuttaCVPR 2024
- Improving the Adversarial Transferability of Vision Transformers with Virtual Dense ConnectionJianping Zhang, Yizhan Huang, Zhuoer Xu, Weibin Wu et al.AAAI 2024 · 22 citations
