Delving into Muon and Beyond: Deep Analysis and Extensions
Xianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He, Rong Xiao
Abstract
The Muon optimizer has recently attracted considerable attention for its strong empirical performance and use of orthogonalized updates on matrix-shaped parameters, yet its underlying mechanisms and relationship to adaptive optimizers such as Adam remain insufficiently understood. In this work, we aim to address these questions through a unified spectral perspective. Specifically, we view Muon as the ( p = 0 ) endpoint of a family of spectral transformations of the form ( U ^p V^ ), and consider additional variants with ( p = 12 ), ( p = 14 ), and ( p = 1 ). These transformations are applied to both first-moment updates, as in momentum SGD, and to root-mean-square (RMS) normalized gradient updates as in Adam. To enable efficient computation, we develop a coupled Newton iteration that avoids explicit singular value decomposition. Across controlled experiments, we find that RMS-normalized updates yield more stable optimization than first-moment updates. Moreover, while spectral compression provides strong stabilization benefits under first-moment updates, the Muon update (( p = 0 )) does not consistently outperform Adam. These results suggest that Muon is best understood as an effective form of spectral normalization, but not a universally superior optimization method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b7ac7a6-dad0-41a9-affb-c0854a9097eaBuilds on5
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 92 citations
- Cautious Optimizers: Improving Training with One Line of CodeKaizhao Liang, Lizhang Chen, Bo Liu, qiang liuICLR 2026 · 38 citations
- LipsFormer: Introducing Lipschitz Continuity to Vision TransformersXianbiao Qi, Jianan Wang, Yihao Chen, Yukai Shi et al.ICLR 2023 · 4 citations
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou et al.ICML 2025
Related papers
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 4 citations
- How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced DataBhavya Vasudeva, Puneesh Deora, Yize Zhao, Vatsal Sharan et al.ICLR 2026 · 16 citations
- Convergence of Muon with Newton-SchulzGyu-Yeol Kim, Min-hwan OhICLR 2026 · 37 citations
- Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix OptimizationZitao Song, Cedar Site Bai, Zhe Zhang, Brian Bullins et al.ICML 2026 · 4 citations
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu et al.ICML 2026 · 7 citations
