Manipulating the Mind's Eye: A-SAGE, the Attention-Based Attack on ViT Explainability
Boshi Zheng, Yan Li, Jiabin Liu
Abstract
The rise of Vision Transformers (ViTs) as cornerstone models in safety-critical applications like autonomous driving and medical diagnosis has shifted the focus from pure accuracy to verifiable trustworthiness. However, the very mechanisms used to explain these models, their internal attention maps, are themselves vulnerable. This creates a critical "trust gap," as the model's apparent reasoning can be maliciously manipulated. To systematically investigate this vulnerability, we introduce A-SAGE (Attention-based Steering Adversarial Generation by Corrupting Explanations), a dual-objective attack framework that forces a model to misclassify an input while simultaneously corrupting its internal attention patterns to generate a misleading explanation. A-SAGE achieves this by optimizing a unified loss that combines a standard classification objective with two explanation-specific terms: an attention entropy loss to diffuse the model's focus and an attention map distortion loss to steer the corrupted explanation towards a desired target. Our primary finding is A-SAGE's exceptional black-box transferability. Using a CaiT-S as a white-box surrogate, adversarial examples generated with imperceptible perturbations achieve attack success rates of 79.4% on ViT-B, 49.7% on ResNet-50, and over 81.5% on other transformers (DeiT-B,TNT-S). Crucially, these successful attacks do not merely destroy the explanation; they generate a coherent but false attention map that deceptively "justifies" the wrong prediction. These results reveal a systemic vulnerability in the core reasoning of modern foundation models, establishing A-SAGE as a critical benchmark for auditing the robustness of AI explainability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a468237-ed4d-45e0-a8a9-b43d3fcfd253Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos et al.ICML 2021 · 1,021 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- On the Robustness of Vision Transformers to Adversarial ExamplesKaleel Mahmood, Rigel Mahmood, Marten van DijkICCV 2021 · 261 citations
Related papers
- Generating Transferable Adversarial Examples against Vision TransformersYuxuan Wang, Jiakai Wang, Zixin Yin, Ruihao Gong et al.ACM MM 2022 · 25 citations
- Harnessing the Computation Redundancy in ViTs to Boost Adversarial TransferabilityJiani Liu, Zhiyuan Wang, Zeliang Zhang, Chao Huang et al.NeurIPS 2025 · 7 citations
- Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch RobustnessGiulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi et al.CVPR 2022 · 33 citations
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt et al.ICCV 2025 · 47 citations
- Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMsChunlong Xie, Kangjie Chen, Shangwei Guo, Shudong Zhang et al.ICML 2026
