On the Convergence of Encoder-only Shallow Transformers
Yongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
摘要
In this paper, we aim to build the global convergence theory of encoder-only shallow Transformers under a realistic setting from the perspective of architectures, initialization, and scaling under a finite width regime. The difficulty lies in how to tackle the softmax in self-attention mechanism, the core ingredient of Transformer. In particular, we diagnose the scaling scheme, carefully tackle the input/output of softmax, and prove that quadratic overparameterization is sufficient for global convergence of our shallow Transformers under commonly-used He/LeCun initialization in practice. Besides, neural tangent kernel (NTK) based analysis is also given, which facilitates a comprehensive comparison. Our theory demonstrates the separation on the importance of different scaling schemes and initialization. We believe our results can pave the way for a better understanding of modern Transformers, particularly on training dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow AnalysisHongru Yang, Bhavya Kailkhura, Zhangyang Wang, Yingbin LiangNeurIPS 2024 · 被引用 14 次
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala 等NeurIPS 2025 · 被引用 10 次
- Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random FeaturesSimone Bombari, Marco MondelliICML 2024 · 被引用 6 次
- Learning to Adapt: In-Context Learning Beyond StationarityZhen Qin, Jiachen Jiang, Zhihui ZhuICLR 2026 · 被引用 1 次
- Transformers Provably Learn Two-Mixture of Linear Classification via Gradient FlowHongru Yang, Zhangyang Wang, Jason D. Lee, Yingbin LiangICLR 2025
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
相关 Paper
- Unraveling the Gradient Descent Dynamics of TransformersBingqing Song, Boran Han, Shuai Zhang, Jie Ding 等NeurIPS 2024 · 被引用 13 次
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari 等NeurIPS 2021 · 被引用 35 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- The Shaped Transformer: Attention Models in the Infinite Depth-and-Width LimitLorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He 等NeurIPS 2023 · 被引用 59 次
- Neural tangent kernels, transportation mappings, and universal approximationZiwei Ji, Matus Telgarsky, Ruicheng XianICLR 2020 · 被引用 45 次
