Memory Efficient Optimizers with 4-bit States
Bingrui Li, Jianfei Chen, Jun Zhu
Abstract
Optimizer states are a major source of memory consumption for training neural networks, limiting the maximum trainable model within given memory budget. Compressing the optimizer states from 32-bit floating points to lower bitwidth is promising to reduce the training memory footprint, while the current lowest achievable bitwidth is 8-bit. In this work, we push optimizer states bitwidth down to 4-bit through a detailed empirical analysis of first and second moments. Specifically, we find that moments have complicated outlier patterns, that current block-wise quantization cannot accurately approximate. We use a smaller block size and propose to utilize both row-wise and column-wise information for better quantization. We further identify a zero point problem of quantizing the second moment, and solve this problem with a linear quantizer that excludes the zero point. Our 4-bit optimizers are evaluated on a wide variety of benchmarks including natural language understanding, machine translation, image classification, and instruction tuning. On all the tasks our optimizers can achieve comparable accuracy with their full-precision counterparts, while enjoying better memory efficiency. * * Code is available at https://github.com/thu-ml/low-bit-optimizers † Corresponding author. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f90833c-12c8-4cf6-b66e-2798e611cd07Cited by top-tier papers34
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 22 citations
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 19 citations
- 4-bit Shampoo for Memory-Efficient Network TrainingSike Wang, Pan Zhou, Jia Li, Hua HuangNeurIPS 2024 · 19 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Achieving low-bit Muon through subspace preservation and grid quantizationHuaijin Wu, Bingrui Li, Yebin Yang, Yi Tu et al.ICLR 2026
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Irrational Complex Rotations Empower Low-bit OptimizersZhen Tian, Xin Zhao, Ji-Rong WenNeurIPS 2025
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 96 citations
- Fine-tuning Quantized Neural Networks with Zeroth-order OptimizationSifeng SHANG, JIAYI ZHOU, Chenyu Lin, Minxian Li et al.ICLR 2026 · 5 citations
