An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Michael Crawshaw, Chirag Modi, Mingrui Liu, Robert Gower
摘要
To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On the Role of Batch Size in Stochastic Conditional Gradient MethodsRustem Islamov, Roman Machacek, Aurelien Lucchi, Antonio Silveti-Falls 等ICML 2026 · 被引用 6 次
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 被引用 4 次
- Step-Size Stability in Stochastic Optimization: A Theoretical PerspectiveFabian Schaipp, Robert Gower, Adrien TaylorICML 2026
- OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit BiasesZixiao Wang, Yifei Shen, Huishuai ZhangICML 2026
它引用的顶会 Paper9
- The Polar Express: Optimal Matrix Sign Methods and their Application to the Muon AlgorithmNoah Amsel, David Persson, Christopher Musco, Robert M. GowerICLR 2026 · 被引用 115 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 被引用 43 次
- Minibatch Stochastic Approximate Proximal Point MethodsHilal Asi, Karan N. Chadha, Gary Cheng, John C. DuchiNeurIPS 2020 · 被引用 22 次
- MoMo: Momentum Models for Adaptive Learning RatesFabian Schaipp, Ruben Ohana, Michael Eickenberg, Aaron Defazio 等ICML 2024 · 被引用 21 次
相关 Paper
- The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural NetworksEitan Gronich, Gal VardiICML 2026
- Error Feedback for Muon and FriendsKaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan, Peter RichtárikICLR 2026 · 被引用 13 次
- Delving into Muon and Beyond: Deep Analysis and ExtensionsXianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He 等ICML 2026 · 被引用 6 次
- FedMuon: Federated Learning with Bias-corrected LMO-based OptimizationYuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. StichICLR 2026 · 被引用 9 次
- Implicit Bias of Spectal Descent and Muon on Multiclass Separable DataChen Fan, Mark Schmidt, Christos ThrampoulidisNeurIPS 2025
