Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement
Yuchen Ren, Zhengyu Zhao, Chenhao Lin, Bo Yang, Lu Zhou, Zhe Liu, Chao Shen
摘要
Vision Transformers (ViTs) have been widely applied in various computer vision and vision-language tasks. To gain insights into their robustness in practical scenarios, transferable adversarial examples on ViTs have been extensively studied. A typical approach to improving adversarial transferability is by refining the surrogate model. However, existing work on ViTs has restricted their surrogate refinement to backward propagation. In this work, we instead focus on Forward Propagation Refinement (FPR) and specifically refine two key modules of ViTs: attention maps and token embeddings. For attention maps, we propose Attention Map Diversification (AMD), which diversifies certain attention maps and also implicitly imposes beneficial gradient vanishing during backward propagation. For token embeddings, we propose Momentum Token Embedding (MTE), which accumulates historical token embeddings to stabilize the forward updates in both the Attention and MLP blocks. We conduct extensive experiments with adversarial examples transferred from ViTs to various CNNs and ViTs, demonstrating that our FPR outperforms the current best (backward) surrogate refinement by up to 7.0% on average. We also validate its superiority against popular defenses and its compatibility with other transfer methods. Codes and appendix are available at https://github.com/RYC-98/FPR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Harnessing the Computation Redundancy in ViTs to Boost Adversarial TransferabilityJiani Liu, Zhiyuan Wang, Zeliang Zhang, Chao Huang 等NeurIPS 2025 · 被引用 7 次
- Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language ModelsYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026 · 被引用 3 次
- Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive InteractionYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026
- Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial TrainingSeongmin Kim, Byung Cheol SongCVPR 2026
它引用的顶会 Paper28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
相关 Paper
- Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token TuningDi Ming, Peng Ren, Yunlong Wang, Xin FengNeurIPS 2024 · 被引用 24 次
- Transferable Adversarial Attack for Both Vision Transformers and Convolutional Networks via Momentum Integrated GradientsWenshuo Ma, Yidong Li, Xiaofeng Jia, Wei XuICCV 2023 · 被引用 62 次
- Random Entangled Tokens for Adversarially Robust Vision TransformerHuihui Gong, Minjing Dong, Siqi Ma, Seyit Camtepe 等CVPR 2024
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong 等ACM MM 2022 · 被引用 25 次
- Transferable Adversarial Attacks on Vision Transformers with Token Gradient RegularizationJianping Zhang, Yizhan Huang, Weibin Wu, Michael R. LyuCVPR 2023
