General Analysis of LMO-based Optimizers: Beyond Bounded Variance
Egor Shulgin, Mohamed Awad, Peter Richtarik, Eduard Gorbunov
Abstract
We study a broad family of momentum Linear Minimization Oracle (LMO) methods that includes normalized SGD with momentum, sign-based (Adam-like) directions, and Muon (spectral) updates. Our focus is on subsampling regimes where the classical uniformly-bounded-variance model can be fragile even for finite-sum objectives on unbounded domains. To obtain subsampling-faithful guarantees, we analyze this LMO family under expected smoothness (ABC condition), which captures common sampling schemes. We establish a unified nonconvex convergence theory via a new self-bounding closure that handles the history-coupling induced by momentum under ABC. Our bounds recover known bounded-variance results as a special case and simplify in strong-growth regimes. Specializing to -nice sampling, we derive explicit batch-size scaling laws, predicting that the optimal momentum must increase with the batch size to maximize sample efficiency. We further identify a theoretical optimal batch size that minimizes total sample complexity. Experiments on linear and matrix regression corroborate these predictions, showing a distinct diagonal shift in the optimal momentum-batch landscape that matches our theoretical scaling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80ec2d6a-8af7-4ece-934a-769c1906706fBuilds on15
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson et al.NeurIPS 2025 · 46 citations
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
Related papers
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 1 citation
- The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural NetworksEitan Gronich, Gal VardiICML 2026
- A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGDRuinan Jin, Xiao Li, Yaoliang Yu, Baoxiang WangICML 2025
- Random Scaling and Momentum for Non-smooth Non-convex OptimizationQinzi Zhang, Ashok CutkoskyICML 2024 · 10 citations
- A Convergence Analysis of Adaptive Optimizers under Floating-point QuantizationXuan Tang, Jichu Li, Difan ZouICLR 2026 · 8 citations
