On the Convergence of Encoder-only Shallow Transformers
Yongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
Abstract
In this paper, we aim to build the global convergence theory of encoder-only shallow Transformers under a realistic setting from the perspective of architectures, initialization, and scaling under a finite width regime. The difficulty lies in how to tackle the softmax in self-attention mechanism, the core ingredient of Transformer. In particular, we diagnose the scaling scheme, carefully tackle the input/output of softmax, and prove that quadratic overparameterization is sufficient for global convergence of our shallow Transformers under commonly-used He/LeCun initialization in practice. Besides, neural tangent kernel (NTK) based analysis is also given, which facilitates a comprehensive comparison. Our theory demonstrates the separation on the importance of different scaling schemes and initialization. We believe our results can pave the way for a better understanding of modern Transformers, particularly on training dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 137484bd-ce89-453b-ab23-a482104e172dCited by top-tier papers5
- Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow AnalysisHongru Yang, Bhavya Kailkhura, Zhangyang Wang, Yingbin LiangNeurIPS 2024 · 14 citations
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala et al.NeurIPS 2025 · 10 citations
- Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random FeaturesSimone Bombari, Marco MondelliICML 2024 · 6 citations
- Learning to Adapt: In-Context Learning Beyond StationarityZhen Qin, Jiachen Jiang, Zhihui ZhuICLR 2026 · 1 citation
- Transformers Provably Learn Two-Mixture of Linear Classification via Gradient FlowHongru Yang, Zhangyang Wang, Jason D. Lee, Yingbin LiangICLR 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
Related papers
- Unraveling the Gradient Descent Dynamics of TransformersBingqing Song, Boran Han, Shuai Zhang, Jie Ding et al.NeurIPS 2024 · 13 citations
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari et al.NeurIPS 2021 · 35 citations
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- The Shaped Transformer: Attention Models in the Infinite Depth-and-Width LimitLorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He et al.NeurIPS 2023 · 59 citations
- Neural tangent kernels, transportation mappings, and universal approximationZiwei Ji, Matus Telgarsky, Ruicheng XianICLR 2020 · 45 citations
