Consensus Is All You Get: The Role of Attention in Transformers
Álvaro Rodríguez Abella, João Pedro Silvestre, Paulo Tabuada
Abstract
A key component of transformers is the attention mechanism orchestrating how each token influences the propagation of every other token along the layers of a transformer. In this paper we provide a rigorous, mathematical analysis of the asymptotic properties of attention in transformers. Although we present several results based on different assumptions, all of them point to the same conclusion, all tokens asymptotically converge to each other, a phenomenon that has been empirically reported in the literature. Our findings are carefully compared with existing theoretical results and illustrated by simulations and experimental studies using the GPT-2 and the GPT-Neo models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0c1634a-9ebc-418e-b571-16cb1ff8424eCited by top-tier papers3
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention DynamicsHugo Koubbi, Louis Hernandez, Matthieu BoussardICML 2026 · 7 citations
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 7 citations
- Attention's forward pass and Frank-WolfeAlbert Alcalde, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026
Builds on6
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao et al.NeurIPS 2022 · 64 citations
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 40 citations
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical SystemsSubhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, Tanmoy ChakrabortyNeurIPS 2021 · 34 citations
Related papers
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 4 citations
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional EncodingDuy-Tung Pham, An Nguyen The, Viet-Hoang Tran, Nhan-Phu Chung et al.NeurIPS 2025 · 2 citations
- Limitations of Normalization in AttentionTimur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu StateNeurIPS 2025
