ZeRO++: Extremely Efficient Collective Communication for Large Model Training
Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Xiaoxia Wu, Connor Holmes, Zhewei Yao, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, Yuxiong He
Abstract
While the Zero Redundancy Optimizer (ZeRO) excels in training large-scale models, it struggles to achieve good throughput in environments with limited bandwidth or small batches where communication becomes a major bottleneck. Inspired by the principles of fine-grained quantization in machine learning algorithms, we designed ZeRO++, an optimizer robust to quantization effects that allows for significant communication volume reduction using low-precision quantization techniques. ZeRO++ composes of three communication volume reduction techniques (low-precision all-gather, data remapping, and low-precision gradient averaging) to significantly reduce the communication volume up to 4x that enables up to 2.16x better throughput at 384 GPU scale. Our results also show ZeRO++ can speedup the RLHF by 3.3x compared to vanilla ZeRO. To verify the convergence of ZeRO++, we test up to 13B model for pretraining with 8/6-bits all gather and up to 30B model for finetuning with 4-bit or 2-bit all gather, and demonstrate on-par accuracy as original ZeRO (aka standard training). As a byproduct, the model trained with ZeRO++ is weight-quantized, which can be directly used for inference without post-training quantization or quantization-aware training. * Equal Contribution. Code has been released as a part of https://github.com/microsoft/DeepSpeed , FY is from University of Houston, LY is from University of Nevada-Reno 37th Conference on Neural Information Processing Systems (NeurIPS 2023). MLsys Workshop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59aaf488-9fc9-46e2-bf9c-9b47200e82d0Cited by top-tier papers6
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin et al.NSDI 2025 · 27 citations
- Revisiting Parameter Server in LLM Post-TrainingXinyi Wan, Penghui Qi, Guangxing Huang, Chaoyi Ruan et al.ICLR 2026 · 2 citations
- DUO: No Compromise to Accuracy DegradationJinda Jia, Cong Xie, Hanlin Lu, Fanjiang Ye et al.NeurIPS 2025
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenXiang Lin, HuangJunTao, LuHan Zhang, Lilaiyi et al.ICML 2026
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
Builds on7
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
Related papers
- Aquavit: Ascending Quantization for Communication-Efficient Vast-Scale Distributed TrainingHong Huang, Jiaxun Ye, Jinhai Yang, Wenjiao Feng et al.KDD 2026
- Fine-tuning Quantized Neural Networks with Zeroth-order OptimizationSifeng SHANG, JIAYI ZHOU, Chenyu Lin, Minxian Li et al.ICLR 2026 · 5 citations
- FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale SpeedSizhe Dang, yangyangGuo, Yanjun Zhao, Xiaodong Zheng et al.ICLR 2026 · 16 citations
- 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence SpeedHanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari et al.ICML 2021 · 106 citations
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang et al.NeurIPS 2024 · 23 citations
