NorMuon: Making Muon more efficient and scalable
Zichong Li, Liming Liu, Chen Liang, Weizhu Chen, Tuo Zhao
摘要
The choice of optimizer significantly impacts the training efficiency and computational costs of large language models (LLMs). Recently, the Muon optimizer has demonstrated promising results by orthogonalizing parameter updates, improving optimization geometry through better conditioning. Despite Muon's emergence as a candidate successor to Adam, the potential for jointly leveraging their strengths-has not been systematically explored. In this work, we bridge this gap by proposing NorMuon (Neuron-wise Normalized Muon), an optimizer that synergistically combines orthogonalization with neuron-level adaptive learning rates. Our analysis reveals that while Muon effectively reduces condition numbers, the resulting updates exhibit highly non-uniform neuron norms, causing certain neurons to dominate the optimization process. NorMuon addresses this imbalance by maintaining second-order momentum statistics for each neuron and applying row-wise normalization after orthogonalization, ensuring balanced parameter utilization while preserving Muon's conditioning benefits. To enable practical deployment at scale, we develop an efficient distributed implementation under the FSDP2 framework that strategically distributes orthogonalization computations across devices. Experiments across multiple model scales demonstrate that NorMuon consistently outperforms both Adam and Muon, achieving 21.74% better training efficiency than Adam and 11.31% improvement over Muon on 1.1B pretraining setting, while maintaining a comparable memory footprint to Muon. Our findings suggest that orthogonalization and adaptive learning rates are complementary rather than competing approaches, opening new avenues for optimizer design in large-scale deep learning. We open-source our implementation at https://github.com/zichongli5/NorMuon.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu 等ICML 2026 · 被引用 7 次
- Can Muon Fine-tune Adam-Pretrained Models?Xingyu Qu, Peigeng Huang, Samuel HorváthICML 2026
它引用的顶会 Paper9
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- How to train data-efficient LLMsNoveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni 等ICLR 2026 · 被引用 106 次
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
相关 Paper
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong 等ICLR 2026 · 被引用 35 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 被引用 9 次
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 被引用 9 次
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language ModelsJungwoo Park, Taewhoo Lee, Chanwoong Yoon, Hyeon Hwang 等ACL 2025 · 被引用 7 次
