On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
Bingrui Li, Wei Huang, Andi Han, Zhanpeng Zhou, Taiji Suzuki, Jun Zhu, Jianfei Chen
Abstract
The Adam optimizer is widely used for transformer optimization in practice, which makes understanding the underlying optimization mechanisms an important problem. However, due to the Adam's complexity, theoretical analysis of how it optimizes transformers remains a challenging task. Fortunately, Sign Gradient Descent (SignGD) serves as an effective surrogate for Adam. Despite its simplicity, theoretical understanding of how SignGD optimizes transformers still lags behind. In this work, we study how SignGD optimizes a two-layer transformer -consisting of a softmax attention layer with trainable query-key parameterization followed by a linear layer -on a linearly separable noisy dataset. We identify four stages in the training dynamics, each exhibiting intriguing behaviors. Based on the training dynamics, we prove the fast convergence but poor generalization of the learned transformer on the noisy dataset. We also show that Adam behaves similarly to SignGD in terms of both optimization and generalization in this setting. Additionally, we find that the poor generalization of SignGD is not solely due to data noise, suggesting that both SignGD and Adam requires high-quality data for real-world tasks. Finally, experiments on synthetic and real-world datasets empirically support our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c06f0884-3ad4-442f-89ff-5611b3d0d52fCited by top-tier papers12
- Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context LearningDake Bu, Wei Huang, Andi Han, Atsushi Nitanda et al.NeurIPS 2024 · 11 citations
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala et al.NeurIPS 2025 · 10 citations
- LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-TuningJunyu Chen, Junzhuo Li, Zhen Peng, Wenjie Wang et al.NeurIPS 2025 · 6 citations
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?Wei Huang, Andi Han, Yujin Song, Yilan Chen et al.NeurIPS 2025 · 4 citations
- Never Saddle for Reparameterized Steepest Descent as Mirror FlowTom Jacobs, Chao Zhou, Rebekka BurkholzICLR 2026 · 3 citations
Builds on46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of NoiseEnea Monzio Compagnoni, Tianlin Liu, Rustem Islamov, Frank Norbert Proske et al.ICLR 2025
- Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effectsKe Liang Xiao, Noah Marshall, Atish Agarwala, Elliot PaquetteICML 2025
- Unraveling the Gradient Descent Dynamics of TransformersBingqing Song, Boran Han, Shuai Zhang, Jie Ding et al.NeurIPS 2024 · 13 citations
- Understanding Adam Requires Better Rotation Dependent AssumptionsTianyue H. Zhang, Lucas Maes, Alan Milligan, Alexia Jolicoeur-Martineau et al.NeurIPS 2025 · 11 citations
- Noise Is Not the Main Factor Behind the Gap Between Sgd and Adam on Transformers, But Sign Descent Might BeFrederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, Mark SchmidtICLR 2023 · 5 citations
