Delving into Muon and Beyond: Deep Analysis and Extensions
Xianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He, Rong Xiao
摘要
The Muon optimizer has recently attracted considerable attention for its strong empirical performance and use of orthogonalized updates on matrix-shaped parameters, yet its underlying mechanisms and relationship to adaptive optimizers such as Adam remain insufficiently understood. In this work, we aim to address these questions through a unified spectral perspective. Specifically, we view Muon as the ( p = 0 ) endpoint of a family of spectral transformations of the form ( U ^p V^ ), and consider additional variants with ( p = 12 ), ( p = 14 ), and ( p = 1 ). These transformations are applied to both first-moment updates, as in momentum SGD, and to root-mean-square (RMS) normalized gradient updates as in Adam. To enable efficient computation, we develop a coupled Newton iteration that avoids explicit singular value decomposition. Across controlled experiments, we find that RMS-normalized updates yield more stable optimization than first-moment updates. Moreover, while spectral compression provides strong stabilization benefits under first-moment updates, the Muon update (( p = 0 )) does not consistently outperform Adam. These results suggest that Muon is best understood as an effective form of spectral normalization, but not a universally superior optimization method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- Cautious Optimizers: Improving Training with One Line of CodeKaizhao Liang, Lizhang Chen, Bo Liu, qiang liuICLR 2026 · 被引用 38 次
- LipsFormer: Introducing Lipschitz Continuity to Vision TransformersXianbiao Qi, Jianan Wang, Yihao Chen, Yukai Shi 等ICLR 2023 · 被引用 4 次
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou 等ICML 2025
相关 Paper
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 被引用 4 次
- How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced DataBhavya Vasudeva, Puneesh Deora, Yize Zhao, Vatsal Sharan 等ICLR 2026 · 被引用 16 次
- Convergence of Muon with Newton-SchulzGyu-Yeol Kim, Min-hwan OhICLR 2026 · 被引用 37 次
- Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix OptimizationZitao Song, Cedar Site Bai, Zhe Zhang, Brian Bullins 等ICML 2026 · 被引用 4 次
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu 等ICML 2026 · 被引用 7 次
