RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
Shenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu, Ruochen Jin, Shuhua Yu, Yaoqing Yang
摘要
Preconditioned adaptive methods have gained significant attention for training deep neural networks, as they capture rich curvature information of the loss landscape . The central challenge in this field lies in balancing preconditioning effectiveness with computational efficiency of implementing the preconditioner. Among recent advances, Muon stands out by using Newton-Schulz iteration to obtain preconditioned updates without explicitly constructing the preconditioning matrix. Despite its advantages, the efficiency of Muon still leaves room for further improvement. In this paper, we introduce RMNP (Row Momentum Normalized Preconditioning), an optimizer that replaces Newton-Schulz iteration with a simple row-wise() normalization operation, motivated by the empirically observed diagonal block structure of the Transformer layerwise Hessian. We empirically verified that orthogonalization and row-wise(on input dim) normalization are asymptotically equivalent in the case of the transformer. This substitution reduces the per-iteration computational complexity from to for an weight matrix while maintaining comparable optimization performance. Theoretically, we establish convergence guarantees for RMNP in the non-convex setting that match recent results for Muon optimizers, achieving the minimax optimal complexity. Extensive experiments on large language model pretraining show that RMNP delivers competitive optimization performance compared with Muon while substantially reducing preconditioning wall-clock time. Our code is available at https://github.com/Dominator-Index/RMNP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li 等NeurIPS 2024 · 被引用 149 次
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- Analytic Insights into Structure and Rank of Neural Network Hessian MapsSidak Pal Singh, Gregor Bachmann, Thomas HofmannNeurIPS 2021 · 被引用 60 次
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- NorMuon: Making Muon more efficient and scalableZichong Li, Liming Liu, Chen Liang, Weizhu Chen 等ICML 2026 · 被引用 57 次
相关 Paper
- The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-NewtonNatalie Abreu, Nikhil Vyas, Sham M. Kakade, Depen MorwaniICLR 2026 · 被引用 21 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- Delving into Muon and Beyond: Deep Analysis and ExtensionsXianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He 等ICML 2026 · 被引用 6 次
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong 等ICLR 2026 · 被引用 35 次
- Convergence of Muon with Newton-SchulzGyu-Yeol Kim, Min-hwan OhICLR 2026 · 被引用 37 次
