Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness
Giulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi, Jan Hendrik Metzen
摘要
Neural architectures based on attention such as vision transformers are revolutionizing image recognition. Their main benefit is that attention allows reasoning about all parts of a scene jointly. In this paper, we show how the global reasoning of (scaled) dot-product attention can be the source of a major vulnerability when confronted with adversarial patch attacks. We provide a theoretical understanding of this vulnerability and relate it to an adversary's ability to misdirect the attention of all queries to a single key token under the control of the adversarial patch. We propose novel adversarial objectives for crafting adversarial patches which target this vulnerability explicitly. We show the effectiveness of the proposed patch attacks on popular image classification (ViTs and DeiTs) and object detection models (DETR). We find that adversarial patches occupying 0.5% of the input can lead to robust accuracies as low as 0% for ViT on ImageNet, and reduce the mAP of DETR on MS COCO to less than 3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Attacking Transformers with Feature Diversity Adversarial PerturbationChenxing Gao, Hang Zhou, Junqing Yu, Yuteng Ye 等AAAI 2024 · 被引用 9 次
- Understanding and Defending Patched-based Adversarial Attacks for Vision TransformerLiang Liu, Yanan Guo, Youtao Zhang, Jun YangICML 2023 · 被引用 7 次
- Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsRunhao Zeng, Xiaoyong Chen, Jiaming Liang, Huisi Wu 等CVPR 2024 · 被引用 6 次
- JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual SteeringRenmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan 等ACM MM 2025 · 被引用 5 次
- Certified Defences Against Adversarial Patch Attacks on Semantic SegmentationMaksym Yatsura, Kaspar Sakmann, N. Grace Hua, Matthias Hein 等ICLR 2023 · 被引用 3 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong 等ACM MM 2022 · 被引用 25 次
- Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?Yonggan Fu, Shunyao Zhang, Shang Wu, Cheng Wan 等ICLR 2022 · 被引用 86 次
- You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?Zenghui Yuan, Pan Zhou, Kai Zou, Yu ChengCVPR 2023
- Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token TuningDi Ming, Peng Ren, Yunlong Wang, Xin FengNeurIPS 2024 · 被引用 24 次
- TrojViT: Trojan Insertion in Vision TransformersMengxin Zheng, Qian Lou, Lei JiangCVPR 2023
