Understanding MARS: When Scaling Momentum Provably Helps
Egor Shulgin, Tamaz Gadaev, Sarit Khirirat, Peter Richtarik
Abstract
MARS (Yuan et al., 2025) has recently emerged as a strong optimizer for large language model (LLM) training by scaling the correction term in momentum-based variance reduction (MVR). However, existing theory does not explain why this modification can improve convergence over the unscaled MVR choice . In this paper, we provide a theoretical explanation for this phenomenon. We introduce -similarity, a refined similarity condition that captures how the scaling coefficient interacts with the stochastic gradient-difference structure. This condition recovers standard similarity at and smoothness at . Using -similarity, we derive convergence guarantees for fixed- MARS whose complexity depends explicitly on and the corresponding -similarity constant. The bound reveals why small values of can be beneficial: they may reduce the similarity term enough to outweigh the penalty from deviating from MVR. We prove that optimizing gives MARS a lower complexity guarantee than MVR. Experiments with MARS-AdamW on GPT-style LLM pretraining corroborate the theory, showing that properly chosen small values of improve token efficiency over and AdamW under a fixed training protocol.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e5cf80a-0395-48c9-8b67-23f60ae54720Builds on12
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- EF21: A New, Simpler, Theoretically Better, and Practically Faster Error FeedbackPeter Richtárik, Igor Sokolov, Ilyas FatkhullinNeurIPS 2021 · 219 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 164 citations
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 32 citations
Related papers
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou et al.ICML 2025
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 1 citation
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 9 citations
- Cautious Optimizers: Improving Training with One Line of CodeKaizhao Liang, Lizhang Chen, Bo Liu, qiang liuICLR 2026 · 38 citations
- Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMsSagnik Mukherjee, Lifan Yuan, Pavan Jayasinha, Dilek Hakkani-Tür et al.ICML 2026 · 4 citations
