ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
Hanwen Cao, Haobo Lu, Xiaosen Wang, Kun He
Abstract
Ensemble-based attacks have been proven to be effective in enhancing adversarial transferability by aggregating the outputs of models with various architectures. However, existing research primarily focuses on refining ensemble weights or optimizing the ensemble path, overlooking the exploration of ensemble models to enhance the transferability of adversarial attacks. To address this gap, we propose applying adversarial augmentation to the surrogate models, aiming to boost overall generalization of ensemble models and reduce the risk of adversarial overfitting. Meanwhile, observing that ensemble Vision Transformers (ViTs) gain less attention, we propose ViT-EnsembleAttack based on the idea of model adversarial augmentation, the first ensemble-based attack method tailored for ViTs to the best of our knowledge. Our approach generates augmented models for each surrogate ViT using three strategies: Multi-head dropping, Attention score scaling, and MLP feature mixing, with the associated parameters optimized by Bayesian optimization. These adversarially augmented models are ensembled to generate adversarial examples. Furthermore, we introduce Automatic Reweighting and Step Size Enlargement modules to boost transferability. Extensive experiments demonstrate that ViT-EnsembleAttack significantly enhances the adversarial transferability of ensemble-based attacks on ViTs, outperforming existing methods by a substantial margin. Code is available at https://github.com/Trustworthy- AI-Group/TransferAttack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bc94fe2-fd27-4f8c-af3a-e2abe2c0b5dfCited by top-tier papers1
Ask how each one uses itBuilds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos et al.ICML 2021 · 1,021 citations
Related papers
- On Improving Adversarial Transferability of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Fahad Shahbaz Khan et al.ICLR 2022 · 111 citations
- An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial TransferabilityBin Chen, Jia-Li Yin, Shukai Chen, Bohao Chen et al.ICCV 2023 · 96 citations
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong et al.ACM MM 2022 · 25 citations
- Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token TuningDi Ming, Peng Ren, Yunlong Wang, Xin FengNeurIPS 2024 · 24 citations
- Boosting Adversarial Transferability via Ensemble Non-AttentionYipeng Zou, Qin Liu, Jie Wu, Yu Peng et al.AAAI 2026
