Lune

EMNLP2025Top-tier venue

Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformers

Jonghyun Hong, Sungyoon Lee

2025Year
2Top-tier citations

Abstract

Attention-based language models commonly rely on the softmax function to convert attention logits into probability distributions. However, this softmax re-weighting can lead to attention entropy collapse, in which attention disproportionately concentrates on a single token, ultimately causing training instability. In this work, we identify the high variance sensitivity of softmax as a primary cause of this collapse. We show that entropy-stable attention methods, which either control or are insensitive to the variance of attention logits, can prevent entropy collapse and enable more stable training. We provide empirical evidence of this effect in both large language models (LLMs) and a small Transformer model composed solely of self-attention and support our findings with theoretical analysis. Moreover, we identify that the concentration of attention probabilities increases the probability matrix norm, leading to the gradient exploding.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 15f6dcab-2c0d-4cf5-8513-5e7cbfab1092

Cited by top-tier papers2

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines