LiMuon: Light and Fast Muon Optimizer for Large Models
Feihu Huang, Yuning Luo, Songcan Chen
Abstract
Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention. More recently, the useful Muon optimizer is specifically designed for matrix-structured parameters of large models. Although some works have begun to study the Muon optimizer, the existing Muon and its variants still suffer from high sample complexity or high memory for large models. To fill this gap, we propose a light and fast Muon (LiMuon) optimizer for training large models, which builds on the momentum-based variance reduced technique and randomized Singular Value Decomposition (SVD). In particular, our LiMuon simultaneously has a lower memory and lower sample complexity than the Muon and its variants. Moreover, we prove that our LiMuon with lower memory has a lower sample complexity of for finding an -stationary solution of non-convex stochastic optimization under the generalized smoothness condition. To further narrow practice and theory gap, we also prove that our LiMuon with Newton-Schulz steps has a lower sample complexity than the Muon with Newton-Schulz steps. Numerical experimental results on training Mamba-130M, Qwen2.5-0.5B and ViT models demonstrate effectiveness of our LiMuon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdbed845-ebf2-4c99-8cc6-03c56a60af31Cited by top-tier papers3
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 37 citations
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network TrainingShenghao Yang, Zhichao Wang, Oleg Balabanov, N. Benjamin Erichson et al.ICML 2026 · 3 citations
- Understanding MARS: When Scaling Momentum Provably HelpsEgor Shulgin, Tamaz Gadaev, Sarit Khirirat, Peter RichtarikICML 2026
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Convergence of Muon with Newton-SchulzGyu-Yeol Kim, Min-hwan OhICLR 2026 · 37 citations
- Training Deep Learning Models with Norm-Constrained LMOsThomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu et al.ICML 2025
Related papers
- MARS: Unleashing the Power of Variance Reduction for Training Large ModelsHuizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou et al.ICML 2025
- Achieving low-bit Muon through subspace preservation and grid quantizationHuaijin Wu, Bingrui Li, Yebin Yang, Yi Tu et al.ICLR 2026
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 17 citations
- Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence GuaranteesThien Hang Nguyen, Huy L. NguyenICML 2025
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 9 citations
