DUO: No Compromise to Accuracy Degradation
Jinda Jia, Cong Xie, Hanlin Lu, Fanjiang Ye, Hao Feng, Daoce Wang, Haibin Lin, Zhi Zhang, Xin Liu
Abstract
Distributed training often suffers from high communication overhead due to large-scale gradient synchronization. Although gradient compression—particularly at 4-bit or even lower precision—significantly reduces transfer volume, it typically results in sacrifice in precision and degradation of the final model accuracy. In this work, we introduce DUO ( Double Update Overlap ), a distributed training framework designed to mitigate accuracy degradation caused by gradient compression without introducing additional overhead. DUO achieves this by inserting an additional high-precision gradient synchronization step into a previously computation-only phase, so that its communication is fully hidden by computation. We provide a comprehensive theoretical proof of convergence for DUO and validate its effectiveness through extensive pre-training experiments on GPT models. Our results indicate that DUO effectively restores accuracy when using 4-bit gradient compression, achieving performance comparable to uncompressed training. Remarkably, DUO maintains minimal accuracy degradation even under extreme compression scenarios, including 1-bit gradients or complete omission of the low-precision gradient communication step (0-bit transmission).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- EF21: A New, Simpler, Theoretically Better, and Practically Faster Error FeedbackPeter Richtárik, Igor Sokolov, Ilyas FatkhullinNeurIPS 2021 · 219 citations
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis et al.ASPLOS 2023 · 64 citations
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen et al.ICML 2023 · 60 citations
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan et al.ASPLOS 2024 · 52 citations
Related papers
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang et al.NeurIPS 2024 · 23 citations
- Step-Ahead Error Feedback for Distributed Training with Compressed GradientAn Xu, Zhouyuan Huo, Heng HuangAAAI 2021 · 17 citations
- Maximizing Communication Efficiency for Large-scale Training via 0/1 AdamYucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa et al.ICLR 2023 · 4 citations
- On Distributed Adaptive Optimization with Gradient CompressionXiaoyun Li, Belhal Karimi, Ping LiICLR 2022 · 34 citations
- Sign bit is enough: a learning synchronization framework for multi-hop all-reduce with ultimate compressionFeijie Wu, Shiqi He, Song Guo, Zhihao Qu et al.DAC 2022 · 3 citations
