Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers
Thiziri Nait Saada, Alireza Naderi, Jared Tanner
Abstract
Attention layers are the core component of transformers, the current state-of-the-art neural network architecture. Alternatives to softmax-based attention are being explored due to its tendency to hinder effective information flow. Even at initialisation, it remains poorly understood why the propagation of signals and gradients through these random networks can be pathological, resulting in issues known as (i) vanishing/exploding gradients and (ii) rank collapse in depth, i.e. when all tokens converge to a single representation along layers. While rank collapse in depth naturally arises from repeated matrix multiplicationsa common pattern across various architectures-we identify an additional and previously unknown challenge unique to softmax attention layers: (iii) rank collapse in width, which occurs as the context length increases. Using Random Matrix Theory, we conduct a rigorous analysis that uncovers a spectral gap between the two largest singular values of the attention matrix as the cause of (iii), which in turn exacerbates (i) and (ii). Building on this insight, we propose a novel yet simple practical solution to mitigate rank collapse in width by removing the outlier eigenvalue(s). Our theoretical framework offers a fresh perspective on recent practical studies, such as [YDX + 24, AGW23], whose ad hoc solutions can now be interpreted as implicit efforts to address the spectral gap issue. This work provides valuable theoretical support for ongoing large-scale empirical research, bringing theory and practice one step closer in the understanding of transformers. 1 * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 525755b9-5d02-4944-98cf-0c7178892e1aCited by top-tier papers9
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for EmbeddingsStephen Zhang, Mustafa Khan, Vardan PapyanNeurIPS 2025 · 18 citations
- When Does Sparsity Mitigate the Curse of Depth in LLMsYao Yao, Xinyuan Song, Sebastian Pokutta, Max Zimmer et al.ICML 2026 · 5 citations
- Universal Redundancies in Time Series Foundation ModelsAnthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William GilpinICML 2026 · 2 citations
- Disentangling Geometry, Performance, and Training in Language ModelsAtharva Kulkarni, Jacob Mitchell Springer, Arjun Subramonian, Swabha SwayamdiptaICML 2026 · 1 citation
- Noise Stability of Transformer ModelsThemistoklis Haris, Zihan Zhang, Yuichi YoshidaICLR 2026
Builds on7
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 212 citations
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Revisiting Over-smoothing in BERT from the Perspective of GraphHan Shi, Jiahui Gao, Hang Xu, Xiaodan Liang et al.ICLR 2022 · 92 citations
Related papers
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2024 · 54 citations
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 15 citations
- Critical attention scaling in long-context transformersShi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe RigolletICLR 2026 · 22 citations
- Spectral Conditioning of Attention Improves Transformer PerformanceHemanth Saratchandran, Simon LuceyNeurIPS 2025 · 9 citations
