Attacking Transformers with Feature Diversity Adversarial Perturbation
Chenxing Gao, Hang Zhou, Junqing Yu, Yuteng Ye, Jiale Cai, Junle Wang, Wei Yang
Abstract
Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit low transferability to other structures and tasks. In this paper, we present a label-free white-box attack approach for ViT-based models that exhibits strong transferability to various black-box models, including most ViT variants, CNNs, and MLPs, even for models developed for other modalities. Our inspiration comes from the feature collapse phenomenon in ViTs, where the critical attention mechanism overly depends on the low-frequency component of features, causing the features in middle-to-end layers to become increasingly similar and eventually collapse. We propose the feature diversity attacker to naturally accelerate this process and achieve remarkable performance and transferability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ae798ee-39c9-4739-8346-b344e5e5c2a8Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
Related papers
- On Improving Adversarial Transferability of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Fahad Shahbaz Khan et al.ICLR 2022 · 111 citations
- Enhancing Transferable Adversarial Attacks on Vision Transformers through Gradient Normalization Scaling and High-Frequency AdaptationZhiyu Zhu, Xinyi Wang, Zhibo Jin, Jiayu Zhang et al.ICLR 2024 · 10 citations
- Improving the Adversarial Transferability of Vision Transformers with Virtual Dense ConnectionJianping Zhang, Yizhan Huang, Zhuoer Xu, Weibin Wu et al.AAAI 2024 · 22 citations
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong et al.ACM MM 2022 · 25 citations
- Transferable Adversarial Attacks on Vision Transformers with Token Gradient RegularizationJianping Zhang, Yizhan Huang, Weibin Wu, Michael R. LyuCVPR 2023
