Convergence of Muon with Newton-Schulz
Gyu-Yeol Kim, Min-hwan Oh
摘要
We analyze Muon as originally proposed and used in practice -- using the momentum orthogonalization with a few Newton-Schulz steps. The prior theoretical results replace this key step in Muon with an exact SVD-based polar factor. We prove that Muon with Newton-Schulz converges to a stationary point at the same rate as the SVD-polar idealization, up to a constant factor for a given number of Newton-Schulz steps. We further analyze this constant factor and prove that it converges to 1 doubly exponentially in and improves with the degree of the polynomial used in Newton-Schulz for approximating the orthogonalization direction. We also prove that Muon removes the typical square-root-of-rank loss compared to its vector-based counterpart, SGD with momentum. Our results explain why Muon with a few low-degree Newton-Schulz steps matches exact-polar (SVD) behavior at a much faster wall-clock time and explain how much momentum matrix orthogonalization via Newton-Schulz benefits over the vector-based optimizer. Overall, our theory justifies the practical Newton-Schulz design of Muon, narrowing its practice-theory gap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LiMuon: Light and Fast Muon Optimizer for Large ModelsFeihu Huang, Yuning Luo, Songcan ChenICML 2026 · 被引用 20 次
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu 等ICML 2026 · 被引用 7 次
- Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided PreconditioningHuan Li, Yiming Dong, Zhouchen LinICML 2026
它引用的顶会 Paper1
相关 Paper
- Delving into Muon and Beyond: Deep Analysis and ExtensionsXianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He 等ICML 2026 · 被引用 6 次
- The Polar Express: Optimal Matrix Sign Methods and their Application to the Muon AlgorithmNoah Amsel, David Persson, Christopher Musco, Robert M. GowerICLR 2026 · 被引用 115 次
- FedMuon: Federated Learning with Bias-corrected LMO-based OptimizationYuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. StichICLR 2026 · 被引用 9 次
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network TrainingShenghao Yang, Zhichao Wang, Oleg Balabanov, N. Benjamin Erichson 等ICML 2026 · 被引用 3 次
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong 等ICLR 2026 · 被引用 35 次
