The emergence of clusters in self-attention dynamics
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe Rigollet
摘要
Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting objects as time tends to infinity. Cluster locations are determined by the initial tokens, confirming context-awareness of representations learned by Transformers. Using techniques from dynamical systems and partial differential equations, we show that the type of limiting object that emerges depends on the spectrum of the value matrix. Additionally, in the one-dimensional case we prove that the self-attention matrix converges to a low-rank Boolean matrix. The combination of these results mathematically confirms the empirical observation made by Vaswani et al. [VSP 17] that leaders appear in a sequence of tokens when processed by Transformers. 1 A classical choice is θ " pW, A, bq P R dˆd ˆRdˆd ˆRd and f θ pxq " W σpAx bq where σ is an elementwise nonlinearity such as the ReLU ([HZRS16b] ).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka 等NeurIPS 2024 · 被引用 54 次
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 被引用 40 次
- How Smooth Is Attention?Valérie Castin, Pierre Ablin, Gabriel PeyréICML 2024 · 被引用 35 次
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 被引用 29 次
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 被引用 24 次
它引用的顶会 Paper5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 被引用 193 次
相关 Paper
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion 等ICML 2026 · 被引用 7 次
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 被引用 7 次
- Consensus Is All You Get: The Role of Attention in TransformersÁlvaro Rodríguez Abella, João Pedro Silvestre, Paulo TabuadaICML 2025
- Emergence of meta-stable clustering in mean-field transformer modelsGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziICLR 2025 · 被引用 2 次
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional EncodingDuy-Tung Pham, An Nguyen The, Viet-Hoang Tran, Nhan-Phu Chung 等NeurIPS 2025 · 被引用 2 次
