Attention's forward pass and Frank-Wolfe
Albert Alcalde, Borjan Geshkovski, Domènec Ruiz-Balet
Abstract
We analyze the hardmax limit of self-attention dynamics for token embeddings in the zero-temperature regime and relate it to finite- behavior. In this limit, the update rule can be viewed as a Frank-Wolfe step for a quadratic objective over the convex hull of the current tokens. When the key-query matrix is negative semidefinite, the dynamics converge with the standard sublinear rate on the quadratic energy, whereas in the positive semidefinite case, extending the hardmax rule to the convex hull induces a Voronoi structure: vertices are stationary, interior points remain in their initial cells, and each token moves along a straight line toward its cell's vertex with exponential convergence under a step-size bounded away from zero. We additionally establish well-posedness of the associated ODE limit in this regime. For finite , we model self-attention as a Markov chain and prove dynamic metastability: interior tokens reach near-vertex configurations in a constant number of steps and remain trapped for times exponential in with high probability, before eventual collapse to some point within the initial convex hull. Thus, hardmax dynamics accurately approximate the finite- process over exponentially long time horizons.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87c8bae1-3d6a-48ee-93b9-b9431c820fd0Cited by top-tier papers4
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention DynamicsHugo Koubbi, Louis Hernandez, Matthieu BoussardICML 2026 · 7 citations
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 7 citations
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
Builds on19
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge et al.ICML 2023 · 153 citations
- Revisiting Over-smoothing in BERT from the Perspective of GraphHan Shi, Jiahui Gao, Hang Xu, Xiaodan Liang et al.ICLR 2022 · 92 citations
Related papers
- Token Sample Complexity of AttentionLéa Bohbot, Cyril Letrouit, Gabriel Peyré, François-Xavier VialardICML 2026 · 1 citation
- Max-Margin Token Selection in Attention MechanismDavoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet OymakNeurIPS 2023 · 67 citations
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 40 citations
- Emergence of meta-stable clustering in mean-field transformer modelsGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziICLR 2025 · 2 citations
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 4 citations
