Emergence of meta-stable clustering in mean-field transformer models
Giuseppe Bruno, Federico Pasqualotto, Andrea Agazzi
摘要
We model the evolution of tokens within a deep stack of Transformer layers as a continuous-time flow on the unit sphere, governed by a mean-field interacting particle system, building on the framework introduced in (Geshkovski et al., 2023) . Studying the corresponding mean-field Partial Differential Equation (PDE), which can be interpreted as a Wasserstein gradient flow, in this paper we provide a mathematical investigation of the long-term behavior of this system, with a particular focus on the emergence and persistence of meta-stable phases and clustering phenomena, key elements in applications like next-token prediction. More specifically, we perform a perturbative analysis of the mean-field PDE around the iid uniform initialization and prove that, in the limit of large number of tokens, the model remains close to a meta-stable manifold of solutions with a given structure (e.g., periodicity). Further, the structure characterizing the meta-stable manifold is explicitly identified, as a function of the inverse temperature parameter of the model, by the index maximizing a certain rescaling of Gegenbauer polynomials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Clustering in Causal Attention MaskingNikita Karagodin, Yury Polyanskiy, Philippe RigolletNeurIPS 2024 · 被引用 40 次
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 被引用 29 次
- Critical attention scaling in long-context transformersShi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe RigolletICLR 2026 · 被引用 22 次
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 被引用 15 次
- Normalization in Attention DynamicsNikita Karagodin, Shu Ge, Yury Polyanskiy, Philippe RigolletNeurIPS 2025 · 被引用 10 次
它引用的顶会 Paper2
- Quantitative Propagation of Chaos for SGD in Wide Neural NetworksValentin De Bortoli, Alain Durmus, Xavier Fontaine, Umut SimsekliNeurIPS 2020 · 被引用 36 次
- Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regimeAndrea Agazzi, Jianfeng LuICLR 2021 · 被引用 3 次
相关 Paper
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 被引用 7 次
- Clustering in Deep Stochastic TransformersLev Fedorov, Michael Sander, Romuald Elie, Pierre Marion 等ICML 2026 · 被引用 7 次
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He 等NeurIPS 2024 · 被引用 10 次
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 被引用 163 次
- Universal Approximation of Mean-Field Models via TransformersShiba Biswal, Karthik Elamvazhuthi, Rishi SonthaliaICML 2025
