Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformers
Jonghyun Hong, Sungyoon Lee
Abstract
Attention-based language models commonly rely on the softmax function to convert attention logits into probability distributions. However, this softmax re-weighting can lead to attention entropy collapse, in which attention disproportionately concentrates on a single token, ultimately causing training instability. In this work, we identify the high variance sensitivity of softmax as a primary cause of this collapse. We show that entropy-stable attention methods, which either control or are insensitive to the variance of attention logits, can prevent entropy collapse and enable more stable training. We provide empirical evidence of this effect in both large language models (LLMs) and a small Transformer model composed solely of self-attention and support our findings with theoretical analysis. Moreover, we identify that the concentration of attention probabilities increases the probability matrix norm, leading to the gradient exploding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15f6dcab-2c0d-4cf5-8513-5e7cbfab1092Cited by top-tier papers2
- Towards Understanding Massive Activations in Attention Sink MechanismHaiyu Wang, Yuanyuan LinICML 2026
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
Related papers
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge et al.ICML 2023 · 153 citations
- Self-attention Networks Localize When QK-eigenspectrum ConcentratesHan Bao, Ryuichiro Hataya, Ryo KarakidaICML 2024 · 16 citations
- Affine-Scaled Attention: Towards Flexible and Stable Transformer AttentionJeongin Bae, baeseong park, Gunho Park, Minsub Kim et al.ICML 2026 · 1 citation
- Self-Adjust SoftmaxChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi et al.EMNLP 2025
- Critical attention scaling in long-context transformersShi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe RigolletICLR 2026 · 22 citations
