Transformers learn through gradual rank increase
Emmanuel Abbe, Samy Bengio, Enric Boix-Adserà, Etai Littwin, Joshua M. Susskind
2023年份
60被引次数
34顶会引用
摘要
We identify incremental learning dynamics in transformers, where the difference between trained and initial weights progressively increases in rank. We rigorously prove this occurs under the simplifying assumptions of diagonal weight matrices and small initialization. Our experiments support the theory and also show that phenomenon can occur in practice without the simplifying assumptions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 被引用 117 次
- Saddle-to-Saddle Dynamics in Diagonal Linear NetworksScott Pesme, Nicolas FlammarionNeurIPS 2023 · 被引用 68 次
- JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and AttentionYuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen 等ICLR 2024 · 被引用 49 次
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 被引用 41 次
- How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation NetworksEtai Littwin, Omid Saremi, Madhu Advani, Vimal Thilak 等NeurIPS 2024 · 被引用 37 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 被引用 119 次
相关 Paper
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training DynamicsZheng-An Chen, Tao LuoNeurIPS 2025 · 被引用 13 次
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 等NeurIPS 2022 · 被引用 161 次
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
- Incremental Learning of Sparse Attention Patterns in TransformersOğuz YükselICML 2026 · 被引用 2 次
- The Implicit Bias of Depth: How Incremental Learning Drives GeneralizationDaniel Gissin, Shai Shalev-Shwartz, Amit DanielyICLR 2020 · 被引用 90 次
