Compression-Aware Gradient Splitting for Collective Communications in Distributed Training
Pranati Majhi, Sabuj Laskar, Abdullah Muzahid, Eun Jung Kim
Abstract
While distributed training is crucial for scaling deep learning models, it incurs significant overhead due to the collective communication of gradients. To alleviate the burden, compression techniques are commonly used to improve network bandwidth utilization. However, compression poses challenges for synchronized AllReduce collective communications, even more so in scalable systems. Non-uniform data sizes resulting from compression can cause bandwidth under-utilization, as faster nodes remain idle while waiting for slower nodes to complete data exchanges, increasing overall communication and consequently, training time. However, the inherent similarity in gradients across consecutive batches presents an opportunity to mitigate these inefficiencies. By leveraging the quantization of gradients and consistent distribution of zeros, the gradients can be partitioned logically to speedup communication. Splitting them into groups with and without zeros can allow different compression approaches for both. The bandwidth under-utilization due to nonuniform data size can also be solved by partitioning the gradients into variable-sized chunks, leading to more balanced compressed data sizes and reduced idle waiting time. We propose two novel strategies in Oscar, where gradient splitting is designed to improve communication and training. Oscar-SW is a novel software-based technique supporting direct AllReduce that splits gradients into probable zeros and non zeros to apply count sketch compression. Oscar-HW, a novel hardware/software codesigned gradient splitting technique is proposed with ASC (Adaptive Stepwise Coding), an encoding technique for gradient compression in distributed training. Oscar-HW dynamically splits fixed-point quantized gradients for AllReduce communications and maximizes bandwidth utilization for state-of-the-art hardware compression techniques. ASC is a variant of Adaptive Arithmetic Coding (AAC) that generates a distinct probability table for each timestep of AllReduce to adapt to its unique value ranges and avoids sending the probability table during the communication of gradients. Our experimental results show that Oscar-SW achievesspeedup and 7 % better accuracy over the SOTA CountSketch algorithm. Oscar-HW achieves an average AllReduce speedup of, and an average end-to-end training speedup of. ASC achieves an average AllReduce speedup ofover Atalanta andover no compression.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 388b0b75-7eef-47c4-96cd-5d9ada2bdaadRelated papers
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho et al.AAAI 2020
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He et al.ICDE 2023 · 9 citations
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui et al.NeurIPS 2020 · 81 citations
- COMPSO: Optimizing Gradient Compression for Distributed Training with Second-Order OptimizersBaixi Sun, Weijin Liu, J. Gregory Pauloski, Jiannan Tian et al.PPoPP 2025 · 8 citations
- DAGC: Data-Aware Adaptive Gradient CompressionRongwei Lu, Jiajun Song, Bin Chen, Laizhong Cui et al.INFOCOM 2023 · 12 citations
