Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
Tsubasa Takahashi, Shojiro Yamabe, Futa Waseda, Kento Sasaki
摘要
Differential Attention (DA) has been proposed as a refinement to standard attention, suppressing redundant or noisy context through a subtractive structure and thereby reducing contextual hallucination. While this design sharpens task-relevant focus, we show that it also introduces a structural fragility under adversarial perturbations. Our theoretical analysis identifies negative gradient alignment—a configuration encouraged by DA’s subtraction—as the key driver of sensitivity amplification, leading to increased gradient norms and elevated local Lipschitz constants. We empirically validate this Fragile Principle through systematic experiments on ViT/DiffViT and evaluations of pretrained CLIP/DiffCLIP, spanning five datasets in total. These results demonstrate higher attack success rates, frequent gradient opposition, and stronger local sensitivity compared to standard attention. Furthermore, depth-dependent experiments reveal a robustness crossover: stacking DA layers attenuates small perturbations via depth-dependent noise cancellation, though this protection fades under larger attack budgets. Overall, our findings uncover a fundamental trade-off: DA improves discriminative focus on clean inputs but increases adversarial vulnerability, underscoring the need to jointly design for selectivity and robustness in future attention mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 被引用 597 次
相关 Paper
- Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language ModelsJun Feng, Shuhong Wu, Hong Sun, Pengfei Zhang 等AAAI 2026
- Pay Less Attention to Function Words for Free Robustness of Vision-Language ModelsQiwei Tian, Chenhao Lin, Zhengyu Zhao, Chao ShenICLR 2026
- Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language ModelsLu Yu, Haiyang Zhang, Changsheng XuNeurIPS 2024 · 被引用 29 次
- Manipulating the Mind's Eye: A-SAGE, the Attention-Based Attack on ViT ExplainabilityBoshi Zheng, Yan Li, Jiabin LiuAAAI 2026
- Towards Transferable Adversarial Attacks on Vision TransformersZhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu 等AAAI 2022 · 被引用 156 次
