On the Variance of the Adaptive Learning Rate and Beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, Jiawei Han
Abstract
The learning rate warmup heuristic achieves remarkable success in stabilizing training, accelerating convergence and improving generalization for adaptive stochastic optimization algorithms like RMSprop and Adam. Here, we study its mechanism in details. Pursuing the theory behind warmup, we identify a problem of the adaptive learning rate (i.e., it has problematically large variance in the early stage), suggest warmup works as a variance reduction technique, and provide both empirical and theoretical evidence to verify our hypothesis. We further propose RAdam, a new variant of Adam, by introducing a term to rectify the variance of the adaptive learning rate. Extensive experimental results on image classification, language modeling, and neural machine translation verify our intuition and demonstrate the effectiveness and robustness of our proposed method. All implementations are available at: this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a661a259-3e43-4628-b6b6-7906cb46f55dCited by top-tier papers299
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- On the Adequacy of Untuned Warmup for Adaptive OptimizationJerry Ma, Denis YaratsAAAI 2021 · 81 citations
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li et al.KDD 2026 · 1 citation
- AdaScale SGD: A User-Friendly Algorithm for Distributed TrainingTyler B. Johnson, Pulkit Agrawal, Haijie Gu, Carlos GuestrinICML 2020 · 41 citations
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou et al.ICML 2025
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural NetworksBohan Wang, Qi Meng, Wei Chen, Tie-Yan LiuICML 2021 · 45 citations
