Always Skip Attention
Yiping Ji, Hemanth Saratchandran, Peyman Moghadam, Simon Lucey
Abstract
We highlight a curious empirical result within modern Vision Transformers (ViTs). Specifically, self-attention catastrophically fails to train unless it is used in conjunction with a skip connection. This is in contrast to other elements of a ViT that continue to exhibit good performance (albeit suboptimal) when skip connections are removed. Further, we show that this critical dependence on skip connections is a relatively new phenomenon, with previous deep architectures (e.g., CNNs) exhibiting good performance in their absence. In this paper, we theoretically characterize that the self-attention mechanism is fundamentally ill-conditioned and is, therefore, uniquely dependent on skip connections for regularization. Additionally, we propose Token Graying (TG), a simple yet effective complement (to skip connections) that further improves the condition of input tokens. We validate our approach in both supervised and self-supervised training methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1121c9e4-a46c-4be9-bf10-538e1c22ea2eCited by top-tier papers5
- Spectral Conditioning of Attention Improves Transformer PerformanceHemanth Saratchandran, Simon LuceyNeurIPS 2025 · 9 citations
- Structured Initialization for Vision TransformersJianqiao Zheng, Xueqian Li, Hemanth Saratchandran, Simon LuceyNeurIPS 2025 · 6 citations
- SineProject: Machine Unlearning for Stable Vision-Language AlignmentArpit Garg, Hemanth Saratchandran, Simon LuceyCVPR 2026 · 2 citations
- Conditioned Initialization for AttentionHemanth Saratchandran, Simon LuceyICLR 2026
- Enhancing Transformers Through Conditioned Embedded TokensHemanth Saratchandran, Simon LuceyICCV 2025
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- Cutting the Skip: Training Residual-Free TransformersYiping Ji, James Martens, Jianqiao Zheng, Ziqin Zhou et al.ICLR 2026 · 7 citations
- Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token TuningDi Ming, Peng Ren, Yunlong Wang, Xin FengNeurIPS 2024 · 24 citations
- Skip-Attention: Improving Vision Transformers by Paying Less AttentionShashanka Venkataramanan, Amir Ghodrati, Yuki M. Asano, Fatih Porikli et al.ICLR 2024 · 42 citations
- A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample ComplexityHongkang Li, Meng Wang, Sijia Liu, Pin-Yu ChenICLR 2023 · 1 citation
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
