SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
Jinda Jia, Cong Xie, Hanlin Lu, Daoce Wang, Hao Feng, Chengming Zhang, Baixi Sun, Haibin Lin, Zhi Zhang, Xin Liu, Dingwen Tao
摘要
Recent years have witnessed a clear trend towards language models with an ever-increasing number of parameters, as well as the growing training overhead and memory usage. Distributed training, particularly through Sharded Data Parallelism (ShardedDP) which partitions optimizer states among workers, has emerged as a crucial technique to mitigate training time and memory usage. Yet, a major challenge in the scalability of ShardedDP is the intensive communication of weights and gradients. While compression techniques can alleviate this issue, they often result in worse accuracy. Driven by this limitation, we propose SDP4Bit (Toward 4Bit Communication Quantization in Sharded Data Parallelism for LLM Training), which effectively reduces the communication of weights and gradients to nearly 4 bits via two novel techniques: quantization on weight differences, and two-level gradient smooth quantization. Furthermore, SDP4Bit presents an algorithm-system co-design with runtime optimization to minimize the computation overhead of compression. In addition to the theoretical guarantees of convergence, we empirically evaluate the accuracy of SDP4Bit on the pre-training of GPT models with up to 6.7 billion parameters, and the results demonstrate a negligible impact on training loss. Furthermore, speed experiments show that SDP4Bit achieves up to 4.08 speedup in end-to-end throughput on a scale of 128 GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- COMPSO: Optimizing Gradient Compression for Distributed Training with Second-Order OptimizersBaixi Sun, Weijin Liu, J. Gregory Pauloski, Jiannan Tian 等PPoPP 2025 · 被引用 8 次
- ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUsJinwu Yang, Jiaan Wu, Zedong Liu, Xinyang Ma 等ISCA 2026 · 被引用 3 次
- STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific DataDaoce Wang, Pascal Grosset, Jesus Pulido, Jiannan Tian 等SC 2025 · 被引用 2 次
- TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM TrainingMan Liu, Xingchen Liu, Xingjian Tian, Bing Lu 等HPDC 2026 · 被引用 1 次
- GPU Travelling: Efficient Confidential Collaborative Training with TEE-Enabled GPUsShixuan Zhao, Zhongshu Gu, Salman Ahmed, Enriquillo Valdez 等CCS 2025
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
相关 Paper
- Quantized Distributed Training of Large Models with Convergence GuaranteesIlia Markov, Adrian Vladu, Qi Guo, Dan AlistarhICML 2023 · 被引用 18 次
- BitDP: Ultra-low-bit Communication for Data Parallelism in LLM TrainingXiaozhe Ren, Qiong LuoAAAI 2026
- DUO: No Compromise to Accuracy DegradationJinda Jia, Cong Xie, Hanlin Lu, Fanjiang Ye 等NeurIPS 2025
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenXiang Lin, HuangJunTao, LuHan Zhang, Lilaiyi 等ICML 2026
- 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence SpeedHanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari 等ICML 2021 · 被引用 106 次
