Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
Atli Kosson, Bettina Messmer, Martin Jaggi
摘要
Learning Rate Warmup is a popular heuristic for training neural networks, especially at larger batch sizes, despite limited understanding of its benefits. Warmup decreases the update size early in training by using lower values for the learning rate . In this work we argue that warmup benefits training by keeping the overall size of limited, counteracting large initial values of . Focusing on small-scale GPT training with AdamW/Lion, we explore the following question: Why and by which criteria are early updates too large? We analyze different metrics for the update size including the -norm, resulting directional change, and impact on the representations of the network, providing a new perspective on warmup. In particular, we find that warmup helps counteract large angular updates as well as a limited critical batch size early in training. Finally, we show that the need for warmup can be significantly reduced or eliminated by modifying the optimizer to explicitly normalize based on the aforementioned metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingShane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray 等NeurIPS 2025 · 被引用 44 次
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 被引用 37 次
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi 等ICLR 2026 · 被引用 11 次
- Scaling with Collapse: Efficient and Predictable Training of LLM FamiliesShane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad 等ICLR 2026 · 被引用 10 次
- Why Do We Need Warm-up? A Theoretical PerspectiveFoivos Alimisis, Rustem Islamov, Aurelien LucchiICML 2026 · 被引用 8 次
它引用的顶会 Paper20
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
相关 Paper
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch TrainingYang Luo, Zangwei Zheng, Ziheng Qin, Zirui Zhu 等ICML 2025
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 被引用 9 次
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 被引用 87 次
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit BiasesZixiao Wang, Yifei Shen, Huishuai ZhangICML 2026
