Understanding Softmax Attention Layers: Exact Mean-Field Analysis on a Toy Problem
Elvis Dohmatob
摘要
Self-attention has emerged as a fundamental component driving the success of modern transformer architectures which power large language models (ChatGPT, Llama, etc.) and various other types of systems. However, a theoretical understanding of how such models actually work is still under active development. The recent work of (Marion et al., 2025) introduced the so-called "single-location regression" problem, which can provably be solved by a simplified self-attention layer but not by linear models, thereby demonstrating a striking functional separation. A rigorous analysis of self-attention with softmax for this problem is challenging due to the coupled nature of the model. In the present work, we use ideas from the classical random energy model in statistical physics to analyze softmax self-attention on the single-location problem. Our analysis yields exact analytic expressions for the population risk in terms of the overlaps between the learned model parameters and those of an oracle. Moreover, we derive a detailed description of the gradient descent dynamics for these overlaps and prove that, under broad conditions, the dynamics converge to the unique oracle attractor. Our work not only advances the understanding of self-attention but also provides key theoretical ideas that are likely to find use in further analyses of even more complex transformer architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
- A Capacity-Based Rationale for Multi-Head AttentionMicah AdlerICML 2026 · 被引用 2 次
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 被引用 394 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- Max-Margin Token Selection in Attention MechanismDavoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet OymakNeurIPS 2023 · 被引用 67 次
- Attention layers provably solve single-location regressionPierre Marion, Raphaël Berthier, Gérard Biau, Claire BoyerICLR 2025
- The Closeness of In-Context Learning and Weight Shifting for Softmax RegressionShuai Li, Zhao Song, Yu Xia, Tong Yu 等NeurIPS 2024 · 被引用 53 次
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 被引用 4 次
- Towards Understanding Transformers in Learning Random WalksWei Shi, Yuan CaoNeurIPS 2025 · 被引用 1 次
