Infinite Limits of Multi-head Transformer Dynamics
Blake Bordelon, Hamza Tahir Chaudhry, Cengiz Pehlevan
摘要
In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime. We identify the set of parameterizations that admit well-defined infinite width and depth limits, allowing the attention layers to update throughout training--a relevant notion of feature learning in these models. We then use tools from dynamical mean field theory (DMFT) to analyze various infinite limits (infinite key/query dimension, infinite heads, and infinite depth) which have different statistical descriptions depending on which infinite limit is taken and how attention layers are scaled. We provide numerical evidence of convergence to the limits and discuss how the parameterization qualitatively influences learned features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li 等NeurIPS 2025 · 被引用 77 次
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 被引用 25 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 被引用 15 次
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 被引用 14 次
相关 Paper
- Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling LimitBlake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin 等ICLR 2024 · 被引用 54 次
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 被引用 6 次
- Transformers Provably Learn Two-Mixture of Linear Classification via Gradient FlowHongru Yang, Zhangyang Wang, Jason D. Lee, Yingbin LiangICLR 2025
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He 等NeurIPS 2024 · 被引用 10 次
- Adaptive kernel predictors from feature-learning infinite limits of neural networksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICML 2025
