BitDP: Ultra-low-bit Communication for Data Parallelism in LLM Training
Xiaozhe Ren, Qiong Luo
摘要
Training large language models (LLMs) with billions of parameters on trillion-token datasets requires distributed data parallelism at increasingly large scales, where gradient synchronization becomes a communication bottleneck, especially in bandwidth-constrained environments. Although gradient quantization presents a promising solution, it faces two key challenges: maintaining training stability and accuracy for transformer architectures and adapting to modern distributed communication systems. In this paper, we propose BitDP, an ultra-low-bit gradient quantization system that reduces communication costs by up to 32× while preserving model accuracy with less than 1% performance degradation. Our approach achieves numerical stability for large transformer models and seamlessly integrates with existing infrastructures. We evaluate BitDP's effectiveness across various LLM sizes, architectures and optimizers. The results demonstrate significant training efficiency improvements while maintaining convergence quality, establishing BitDP as a scalable and reliable solution for real-world LLM training at industrial scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang 等NeurIPS 2024 · 被引用 23 次
相关 Paper
- Quantized Distributed Training of Large Models with Convergence GuaranteesIlia Markov, Adrian Vladu, Qi Guo, Dan AlistarhICML 2023 · 被引用 18 次
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenXiang Lin, HuangJunTao, LuHan Zhang, Lilaiyi 等ICML 2026
- Low-Bit Quantization Favors Undertrained LLMsXu Ouyang, Tao Ge, Thomas Hartvigsen, Zhisong Zhang 等ACL 2025 · 被引用 3 次
- BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-DistillationDayou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo 等ACL 2024 · 被引用 16 次
- Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model ParallelismSameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo 等NeurIPS 2025 · 被引用 12 次
