The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris J. Maddison, Dan Roy
摘要
In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softmax-based attention model with skip connections in the proportional limit of infinite-depth-and-width. We show that at initialization the limiting distribution can be described by a stochastic differential equation (SDE) indexed by the depth-to-width ratio. To achieve a well-defined stochastic limit, the Transformer's attention mechanism is modified by centering the Softmax output at identity, and scaling the Softmax logits by a width-dependent temperature parameter. We examine the stability of the network through the corresponding SDE, showing how the scale of both the drift and diffusion can be elegantly controlled with the aid of residual connections. The existence of a stable SDE implies that the covariance structure is well-behaved, even for very large depth and width, thus preventing the notorious issues of rank degeneracy in deep attention models. Finally, we show, through simulations, that the SDE provides a surprisingly good description of the corresponding finite-size model. We coin the name shaped Transformer for these architectural modifications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li 等NeurIPS 2025 · 被引用 77 次
- Tensor Programs VI: Feature Learning in Infinite Depth Neural NetworksGreg Yang, Dingli Yu, Chen Zhu, Soufiane HayouICLR 2024 · 被引用 77 次
- The Impact of Initialization on LoRA Finetuning DynamicsSoufiane Hayou, Nikhil Ghosh, Bin YuNeurIPS 2024 · 被引用 63 次
- Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling LimitBlake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin 等ICLR 2024 · 被引用 54 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 被引用 217 次
相关 Paper
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 被引用 51 次
- Width and Depth Limits Commute in Residual NetworksSoufiane Hayou, Greg YangICML 2023 · 被引用 23 次
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 被引用 4 次
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 被引用 15 次
- On the Convergence of Encoder-only Shallow TransformersYongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2023 · 被引用 17 次
