Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token Tuning
Di Ming, Peng Ren, Yunlong Wang, Xin Feng
Abstract
Vision transformers (ViTs) perform exceptionally well in various computer vision tasks but remain vulnerable to adversarial attacks. Recent studies have shown that the transferability of adversarial examples exists for CNNs, and the same holds true for ViTs. However, existing ViT attacks aggressively regularize the largest token gradients to exact zero within each layer of the surrogate model, overlooking the interactions between layers, which limits their transferability in attacking black-box models. Therefore, in this paper, we focus on boosting the transferability of adversarial attacks on ViTs through adaptive token tuning (ATT). Specifically, we propose three optimization strategies: an adaptive gradient re-scaling strategy to reduce the overall variance of token gradients, a self-paced patch out strategy to enhance the diversity of input tokens, and a hybrid token gradient truncation strategy to weaken the effectiveness of attention mechanism. We demonstrate that scaling correction of gradient changes using gradient variance across different layers can produce highly transferable adversarial examples. In addition, introducing attentional truncation can mitigate the overfitting over complex interactions between tokens in deep ViT layers to further improve the transferability. On the other hand, using feature importance as a guidance to discard a subset of perturbation patches in each iteration, along with combining self-paced learning and progressively more sampled attacks, significantly enhances the transferability over attacks that use all perturbation patches. Extensive experiments conducted on ViTs, undefended CNNs, and defended CNNs validate the superiority of our proposed ATT attack method. On average, our approach improves the attack performance by 10.1% compared to state-of-the-art transfer-based attacks. Notably, we achieve the best attack performance with an average of 58.3% on three defended CNNs. Code is available at https:/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fda196ee-8d75-4336-a1bd-fabefe21da21Cited by top-tier papers6
- Guiding a Diffusion Model by Swapping Its TokensWeijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran et al.CVPR 2026 · 2 citations
- Boosting Generative Adversarial Transferability with Self-Supervised Vision Transformer FeaturesShangbo Wu, Yu-an Tan, Ruinan Ma, Wencong Ma et al.ICCV 2025
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu et al.USENIX Security 2026
- MADA-Attack: Transferable Multi-modal Attention Distraction Adversarial Attack against Vision Language ModelsZhihan Qin, Jiahao Chen, Chunyi Zhou, Yuwen Pu et al.ICML 2026
- Low-Rank and Sparsity Are All You Need: Exploring Robust Hierarchical Latent Subspaces for Transferable Adversarial AttackShuangshuang Pu, Wen Yang, Min Li, guodong liu et al.ICML 2026
Builds on39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- On Improving Adversarial Transferability of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Fahad Shahbaz Khan et al.ICLR 2022 · 111 citations
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu et al.AAAI 2022 · 156 citations
- Enhancing Transferable Adversarial Attacks on Vision Transformers through Gradient Normalization Scaling and High-Frequency AdaptationZhiyu Zhu, Xinyi Wang, Zhibo Jin, Jiayu Zhang et al.ICLR 2024 · 10 citations
- Transferable Adversarial Attacks on Vision Transformers with Token Gradient RegularizationJianping Zhang, Yizhan Huang, Weibin Wu, Michael R. LyuCVPR 2023
- Improving the Adversarial Transferability of Vision Transformers with Virtual Dense ConnectionJianping Zhang, Yizhan Huang, Zhuoer Xu, Weibin Wu et al.AAAI 2024 · 22 citations
