Critical attention scaling in long-context transformers
Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet
Abstract
As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse. While effectively addresses this deficiency by rescaling attention scores with a polylogarithmic factor , theoretical justification for this approach remains lacking.
We analyze a simplified yet tractable model that magnifies the effect of attention scaling. In this model, attention exhibits a phase transition governed by the scaling factor : insufficient scaling collapses all tokens to a single direction, while excessive scaling reduces attention to identity, thereby eliminating meaningful interactions between tokens. Our main result identifies the critical scaling and provides a rigorous justification for attention scaling in YaRN and Qwen, clarifying why logarithmic scaling maintains sparse, content-adaptive attention at large context lengths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da6964f-1f39-40c2-b9db-ef04846c1e3dCited by top-tier papers5
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention DynamicsHugo Koubbi, Louis Hernandez, Matthieu BoussardICML 2026 · 7 citations
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 7 citations
- Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language ModelingXingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu et al.ACL 2026 · 3 citations
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
Builds on8
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 40 citations
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
Related papers
- Token Sample Complexity of AttentionLéa Bohbot, Cyril Letrouit, Gabriel Peyré, François-Xavier VialardICML 2026 · 1 citation
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- Self-attention Networks Localize When QK-eigenspectrum ConcentratesHan Bao, Ryuichiro Hataya, Ryo KarakidaICML 2024 · 16 citations
- Length-Induced Embedding Collapse in PLM-based ModelsYuqi Zhou, Sunhao Dai, Zhanshuo Cao, Xiao Zhang et al.ACL 2025 · 8 citations
