Understanding Softmax Attention Layers: Exact Mean-Field Analysis on a Toy Problem
Elvis Dohmatob
Abstract
Self-attention has emerged as a fundamental component driving the success of modern transformer architectures which power large language models (ChatGPT, Llama, etc.) and various other types of systems. However, a theoretical understanding of how such models actually work is still under active development. The recent work of (Marion et al., 2025) introduced the so-called "single-location regression" problem, which can provably be solved by a simplified self-attention layer but not by linear models, thereby demonstrating a striking functional separation. A rigorous analysis of self-attention with softmax for this problem is challenging due to the coupled nature of the model. In the present work, we use ideas from the classical random energy model in statistical physics to analyze softmax self-attention on the single-location problem. Our analysis yields exact analytic expressions for the population risk in terms of the overlaps between the learned model parameters and those of an oracle. Moreover, we derive a detailed description of the gradient descent dynamics for these overlaps and prove that, under broad conditions, the dynamics converge to the unique oracle attractor. Our work not only advances the understanding of self-attention but also provides key theoretical ideas that are likely to find use in further analyses of even more complex transformer architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42a61ab8-db92-4bd1-beed-aa86fd059c44Cited by top-tier papers2
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro et al.ICLR 2026 · 7 citations
- A Capacity-Based Rationale for Multi-Head AttentionMicah AdlerICML 2026 · 2 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 324 citations
Related papers
- Max-Margin Token Selection in Attention MechanismDavoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet OymakNeurIPS 2023 · 67 citations
- Attention layers provably solve single-location regressionPierre Marion, Raphaël Berthier, Gérard Biau, Claire BoyerICLR 2025
- The Closeness of In-Context Learning and Weight Shifting for Softmax RegressionShuai Li, Zhao Song, Yu Xia, Tong Yu et al.NeurIPS 2024 · 53 citations
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 4 citations
- Towards Understanding Transformers in Learning Random WalksWei Shi, Yuan CaoNeurIPS 2025 · 1 citation
