Perceptrons and Localization of Attention’s Mean-Field Landscape
Antonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-Balet
Abstract
The forward pass of a Transformer can be seen as an interacting particle system on the unit sphere: time plays the role of layers, particles that of token embeddings, and the unit sphere idealizes layer normalization. In some weight settings the system can even be seen as a gradient flow for an explicit energy, and one can make sense of the infinite context length (mean-field) limit thanks to Wasserstein gradient flows. In this paper we study the effect of the perceptron block in this setting, and show that critical points are generically atomic and localized on subsets of the sphere.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d52d87b-3568-43fa-a641-fd9b073e6f63Builds on12
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 40 citations
- A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product AttentionHugo Cui, Freya Behrens, Florent Krzakala, Lenka ZdeborováNeurIPS 2024 · 35 citations
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical SystemsSubhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, Tanmoy ChakrabortyNeurIPS 2021 · 34 citations
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
Related papers
- Emergence of meta-stable clustering in mean-field transformer modelsGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziICLR 2025 · 2 citations
- Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention LandscapeJuno Kim, Taiji SuzukiICML 2024 · 42 citations
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion et al.ICML 2026 · 7 citations
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He et al.NeurIPS 2024 · 10 citations
- Towards Understanding Inductive Bias in Transformers: A View From InfinityItay Lavie, Guy Gur-Ari, Zohar RingelICML 2024 · 11 citations
