Achieving low-bit Muon through subspace preservation and grid quantization
Huaijin Wu, Bingrui Li, Yebin Yang, Yi Tu, Zhanpeng Zhou, Jianfei Chen, Junchi Yan
摘要
Training Large Language Models (LLMs) faces severe memory constraints due to the increasing size of model parameters and optimizer states. The Muon optimizer, which is based on matrix orthogonalization, has recently demonstrated significant potential and offers considerable memory advantages over AdamW by utilizing only the first moment. However, how to apply memory-reduction techniques to further compress the optimizer states of Muon remains underexplored. Directly applying existing methods may encounter significant difficulties due to the orthogonalization process. In this work, we investigate the low-bit compression of Muon and systematically analyze the quantization error exacerbated by orthogonalization. We identify that the error primarily originates from the top singular subspace and the outlier patterns of moment matrix appearing across both dimensions. To address this, we propose 4-bit-Muon-GRASP (GRid And Subspace Preserving), which compresses the Muon optimizer states to 4 bits using grid quantization, while preserving the top singular subspace with minimal overhead. We evaluate 4-bit-Muon-GRASP through pre-training on LLaMA-130M, 350M, and 1.1B architectures and fine-tuning on 7B models for various reasoning tasks. Extensive experiment results show that our 4-bit-Muon-GRASP achieves accuracy comparable to full-precision counterparts while reducing training memory consumption by up to 28%. The source code is publicly available at https://github.com/wuhuaijin/lowbit-Muon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Memory Efficient Optimizers with 4-bit StatesBingrui Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 72 次
- A Convergence Analysis of Adaptive Optimizers under Floating-point QuantizationXuan Tang, Jichu Li, Difan ZouICLR 2026 · 被引用 8 次
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenXiang Lin, HuangJunTao, LuHan Zhang, Lilaiyi 等ICML 2026
- Taming Momentum: Rethinking Optimizer States Through Low-Rank ApproximationZhengbo Wang, Jian Liang, Ran He, Zilei Wang 等ICLR 2026 · 被引用 3 次
- Fine-tuning Quantized Neural Networks with Zeroth-order OptimizationSifeng SHANG, JIAYI ZHOU, Chenyu Lin, Minxian Li 等ICLR 2026 · 被引用 5 次
