Taming Transformer Without Using Learning Rate Warmup
Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou, Rong Xiao
Abstract
Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and using an obviously lower learning rate is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal the rationale behind the model crash phenomenon in the training process, termed spectral energy concentration of W q ⊤ W k , which is the reason for a malignant entropy collapse, where W q and W k are the projection matrices for the query and the key in Transformer, respectively. To remedy this problem, motivated by Weyl's Inequality, we present a novel optimization strategy, i.e., making the weight updating in successive steps smooth-if the ratio ) is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of σ 1 (W t -1 ) σ 1 (∇W t ) , where ∇W t is the updating quantity in step t . Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these Transformers without using learning rate warmup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27767bd8-27a6-4586-80a0-0107609eeb72Cited by top-tier papers4
- SimpleGPT: Improving GPT via A Simple Normalization StrategyMarco Chen, Xianbiao Qi, Yelin He, Jiaquan Ye et al.ICML 2026 · 2 citations
- DNT: a Deeply Normalized Transformer that can be trained by Momentum SGDXianbiao Qi, Marco Chen, Wenjie Xiao, Jiaquan Ye et al.ICLR 2026 · 1 citation
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 1 citation
- Conditioned Initialization for AttentionHemanth Saratchandran, Simon LuceyICLR 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge et al.ICML 2023 · 153 citations
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 212 citations
- LipsFormer: Introducing Lipschitz Continuity to Vision TransformersXianbiao Qi, Jianan Wang, Yihao Chen, Yukai Shi et al.ICLR 2023 · 4 citations
- Initialization of Large Language Models via Reparameterization to Mitigate Loss SpikesKosuke Nishida, Kyosuke Nishida, Kuniko SaitoEMNLP 2024 · 2 citations
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
