Fine-grained Local Sensitivity Analysis of Standard Dot-Product Self-Attention
Aaron J. Havens, Alexandre Araujo, Huan Zhang, Bin Hu
Abstract
Self-attention has been widely used in various machine learning models, such as vision transformers. The standard dot-product self-attention is arguably the most popular structure, and there is a growing interest in understanding the mathematical properties of such attention mechanisms. This paper presents a fine-grained local sensitivity analysis of the standard dot-product selfattention, leading to new non-vacuous certified robustness results for vision transformers. Despite the well-known fact that dot-product selfattention is not (globally) Lipschitz, we develop new theoretical analysis of Local Fine-grained Attention Sensitivity (LoFAST) quantifying the effect of input feature perturbations on the attention output. Our analysis reveals that the local sensitivity of dot-product self-attention to ℓ 2 perturbations can actually be controlled by several key quantities associated with the attention weight matrices and the unperturbed input. We empirically validate our theoretical findings by computing non-vacuous certified ℓ 2 -robustness for vision transformers on CIFAR-10 and SVHN datasets. The code for LoFAST is available at https: //github.com/AaronHavens/LoFAST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Certified Robustness Under Bounded Levenshtein DistanceElías Abad-Rocamora, Grigorios Chrysos, Volkan CevherICLR 2025
- Efficient Robust Conformal Prediction via Lipschitz-Bounded NetworksThomas Massena, Léo Andéol, Thibaut Boissin, Franck Mamalet et al.ICML 2025
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
Related papers
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
- How Smooth Is Attention?Valérie Castin, Pierre Ablin, Gabriel PeyréICML 2024 · 35 citations
- Alias-Free ViT: Fractional Shift Invariance via Linear AttentionHagay Michaeli, Daniel SoudryNeurIPS 2025 · 2 citations
- Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch RobustnessGiulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi et al.CVPR 2022 · 33 citations
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models Against Gaussian NoiseBum Jun Kim, Makoto Kawano, Yusuke Iwasawa, Yutaka MatsuoICML 2026
