On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, Sanjeev Arora
Abstract
Approximating Stochastic Gradient Descent (SGD) as a Stochastic Differential Equation (SDE) has allowed researchers to enjoy the benefits of studying a continuous optimization trajectory while carefully preserving the stochasticity of SGD. Analogous study of adaptive gradient methods, such as RMSprop and Adam, has been challenging because there were no rigorously proven SDE approximations for these methods. This paper derives the SDE approximations for RMSprop and Adam, giving theoretical guarantees of their correctness as well as experimental validation of their applicability to common large-scaling vision and language settings. A key practical result is the derivation of a to adjust the optimization hyperparameters of RMSprop and Adam when changing batch size, and its empirical validation in deep learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers44
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25, 000 SubjectsYassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia et al.NeurIPS 2025 · 106 citations
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 76 citations
- Implicit Bias of AdamW: ℓ∞-Norm Constrained OptimizationShuo Xie, Zhiyuan LiICML 2024 · 46 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
Related papers
- Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of NoiseEnea Monzio Compagnoni, Tianlin Liu, Rustem Islamov, Frank Norbert Proske et al.ICLR 2025
- Surge Phenomenon in Optimal Learning Rate and Batch Size ScalingShuaipeng Li, Penghao Zhao, Hailin Zhang, Xingwu Sun et al.NeurIPS 2024 · 33 citations
- On the Implicit Bias of AdamMatias D. Cattaneo, Jason M. Klusowski, Boris ShigidaICML 2024 · 26 citations
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima et al.NeurIPS 2024 · 32 citations
- Non-asymptotic Analysis of Biased Adaptive Stochastic ApproximationSobihan Surendran, Adeline Fermanian, Antoine Godichon-Baggioni, Sylvain Le CorffNeurIPS 2024 · 7 citations
