Dimension-adapted Momentum Outscales SGD
Damien Ferbach, Katie Everett, Gauthier Gidel, Elliot Paquette, Courtney Paquette
摘要
We investigate scaling laws for stochastic momentum algorithms with small batch on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear RegressionTingkai Yan, Haodong Wen, Binghui Li, Kairong Luo 等ICLR 2026 · 被引用 12 次
- High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizesAukosh Jagannath, Taj Jones-McCormick, Varnan SarangianICLR 2026 · 被引用 1 次
- Predicting Large Model Test Losses with a Noisy Quadratic SystemChuning Li, Chris MaddisonICML 2026
- Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?Jihwan Kim, Dogyoon Song, Chulhee YunICLR 2026
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- The Road Less ScheduledAaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko 等NeurIPS 2024 · 被引用 208 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
相关 Paper
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic ModelsCourtney Paquette, Elliot PaquetteNeurIPS 2021 · 被引用 20 次
- Accelerating SGD with momentum for over-parameterized learningChaoyue Liu, Mikhail BelkinICLR 2020 · 被引用 93 次
- Nesterov acceleration in benignly non-convex landscapesKanan Gupta, Stephan WojtowytschICLR 2025
- Momentum Further Constrains Sharpness at the Edge of Stochastic StabilityArseniy Andreyev, Advikar Ananthkumar, Marc Walden, Tomaso A Poggio 等ICML 2026 · 被引用 5 次
- Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High DimensionsKiwon Lee, Andrew N. Cheng, Elliot Paquette, Courtney PaquetteNeurIPS 2022 · 被引用 22 次
