Non-asymptotic Convergence of Training Transformers for Next-token Prediction
Ruiquan Huang, Yingbin Liang, Jing Yang
Abstract
Transformers have achieved extraordinary success in modern machine learning due to their excellent ability to handle sequential data, especially in next-token prediction (NTP) tasks. However, the theoretical understanding of their performance in NTP is limited, with existing studies focusing mainly on asymptotic performance. This paper provides a fine-grained non-asymptotic analysis of the training dynamics of a one-layer transformer consisting of a self-attention module followed by a feed-forward layer. We first characterize the essential structural properties of training datasets for NTP using a mathematical framework based on partial orders. Then, we design a two-stage training algorithm, where the pre-processing stage for training the feed-forward layer and the main stage for training the attention layer exhibit fast convergence performance. Specifically, both layers converge sub-linearly to the direction of their corresponding max-margin solutions. We also show that the cross-entropy loss enjoys a linear convergence rate. Furthermore, we show that the trained transformer presents non-trivial prediction ability with dataset shift, which sheds light on the remarkable generalization performance of transformers. Our analysis technique involves the development of novel properties on the attention gradient and further in-depth analysis of how these properties contribute to the convergence of the training process. Our experiments further validate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e128080-fdc9-451f-b6ff-55da21d877c6Cited by top-tier papers9
- How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic InterpretabilityShawn Im, Changdae Oh, Zhen Fang, Sharon LiICLR 2026 · 4 citations
- Unlabeled Data Can Provably Enhance In-Context Learning of TransformersRenpu Liu, Jing YangNeurIPS 2025 · 3 citations
- Decoupling Positional and Symbolic Attention in TransformersFelipe Urrutia, Jorge Salas, Alexander Kozachinskiy, Cristian Buc Calderon et al.ICLR 2026 · 3 citations
- In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separationFrank Cole, Yuxuan Zhao, Yulong Lu, Tianhao ZhangNeurIPS 2025 · 1 citation
- The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-trainingHongtao Zhang, WenJie Zhou, Chenxi Jia, Wei Chen et al.ICML 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 324 citations
Related papers
- In-context Convergence of TransformersYu Huang, Yuan Cheng, Yingbin LiangICML 2024 · 114 citations
- On the Generalization Ability of Next-Token-Prediction PretrainingZhihao Li, Xue Jiang, Liyuan Liu, Xuelin Zhang et al.ICML 2025
- Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer TransformerYuandong Tian, Yiping Wang, Beidi Chen, Simon S. DuNeurIPS 2023 · 125 citations
- Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow AnalysisHongru Yang, Bhavya Kailkhura, Zhangyang Wang, Yingbin LiangNeurIPS 2024 · 14 citations
- Implicit Optimization Bias of Next-token Prediction in Linear ModelsChristos ThrampoulidisNeurIPS 2024 · 19 citations
