From Muon to Gluon: Bridging Theory and Practice of LMO-based Optimizers for LLMs
Artem Riabinin, Egor Shulgin, Kaja Gruntkowska, Peter Richtarik
摘要
Recent developments in deep learning optimization have brought about radically new algorithms based on the Linear Minimization Oracle (LMO) framework, such as Muon [12] and Scion [24] . After over a decade of Adam's dominance, these LMO-based methods are emerging as viable replacements, offering several practical advantages such as improved memory efficiency, better hyperparameter transferability, and most importantly, superior empirical performance on large-scale tasks, including LLM training. However, a significant gap remains between their practical use and our current theoretical understanding: prior analyses (1) overlook the layer-wise LMO application of these optimizers in practice, and (2) rely on an unrealistic smoothness assumption, leading to impractically small stepsizes. To address both, we propose a new LMObased method called Gluon, capturing prior theoretically analyzed methods as special cases, and introduce a new refined generalized smoothness model that captures the layer-wise geometry of neural networks, matches the layer-wise practical implementation of Muon and Scion, and leads to convergence guarantees with strong practical predictive power. Unlike prior results, our theoretical stepsizes closely match the fine-tuned values reported by Pethick et al. [24]. Our experiments with NanoGPT and CNN confirm that our assumption holds along the optimization trajectory, ultimately closing the gap between theory and practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SGD with Adaptive Preconditioning: Unified Analysis and Momentum AccelerationDmitry KovalevICLR 2026 · 被引用 13 次
- Extragradient Method for -Lipschitz Root-finding ProblemsSayantan Choudhury, Nicolas LoizouNeurIPS 2025 · 被引用 5 次
- General Analysis of LMO-based Optimizers: Beyond Bounded VarianceEgor Shulgin, Mohamed Awad, Peter Richtarik, Eduard GorbunovICML 2026
它引用的顶会 Paper9
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang 等NeurIPS 2022 · 被引用 111 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- A Guide Through the Zoo of Biased SGDYury Demidovich, Grigory Malinovsky, Igor Sokolov, Peter RichtárikNeurIPS 2023 · 被引用 56 次
- Understanding the Generalization of Adam in Learning Neural Networks with Proper RegularizationDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICLR 2023 · 被引用 6 次
相关 Paper
- Error Feedback for Muon and FriendsKaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan, Peter RichtárikICLR 2026 · 被引用 13 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- Training Deep Learning Models with Norm-Constrained LMOsThomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu 等ICML 2025
- Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity HandlingDmitrii Feoktistov, Timofey Belinsky, Andrey Veprikov, Amir Zainullin 等ICML 2026
- FedMuon: Federated Learning with Bias-corrected LMO-based OptimizationYuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. StichICLR 2026 · 被引用 9 次
