Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformers
Jonghyun Hong, Sungyoon Lee
摘要
Attention-based language models commonly rely on the softmax function to convert attention logits into probability distributions. However, this softmax re-weighting can lead to attention entropy collapse, in which attention disproportionately concentrates on a single token, ultimately causing training instability. In this work, we identify the high variance sensitivity of softmax as a primary cause of this collapse. We show that entropy-stable attention methods, which either control or are insensitive to the variance of attention logits, can prevent entropy collapse and enable more stable training. We provide empirical evidence of this effect in both large language models (LLMs) and a small Transformer model composed solely of self-attention and support our findings with theoretical analysis. Moreover, we identify that the concentration of attention probabilities increases the probability matrix norm, leading to the gradient exploding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Understanding Massive Activations in Attention Sink MechanismHaiyu Wang, Yuanyuan LinICML 2026
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
相关 Paper
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge 等ICML 2023 · 被引用 153 次
- Self-attention Networks Localize When QK-eigenspectrum ConcentratesHan Bao, Ryuichiro Hataya, Ryo KarakidaICML 2024 · 被引用 16 次
- Affine-Scaled Attention: Towards Flexible and Stable Transformer AttentionJeongin Bae, baeseong park, Gunho Park, Minsub Kim 等ICML 2026 · 被引用 1 次
- Self-Adjust SoftmaxChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi 等EMNLP 2025
- Critical attention scaling in long-context transformersShi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe RigolletICLR 2026 · 被引用 22 次
