Lune

HPCA2026Top-tier venue

Compression-Aware Gradient Splitting for Collective Communications in Distributed Training

Pranati Majhi, Sabuj Laskar, Abdullah Muzahid, Eun Jung Kim

2026Year

Abstract

While distributed training is crucial for scaling deep learning models, it incurs significant overhead due to the collective communication of gradients. To alleviate the burden, compression techniques are commonly used to improve network bandwidth utilization. However, compression poses challenges for synchronized AllReduce collective communications, even more so in scalable systems. Non-uniform data sizes resulting from compression can cause bandwidth under-utilization, as faster nodes remain idle while waiting for slower nodes to complete data exchanges, increasing overall communication and consequently, training time. However, the inherent similarity in gradients across consecutive batches presents an opportunity to mitigate these inefficiencies. By leveraging the quantization of gradients and consistent distribution of zeros, the gradients can be partitioned logically to speedup communication. Splitting them into groups with and without zeros can allow different compression approaches for both. The bandwidth under-utilization due to nonuniform data size can also be solved by partitioning the gradients into variable-sized chunks, leading to more balanced compressed data sizes and reduced idle waiting time. We propose two novel strategies in Oscar, where gradient splitting is designed to improve communication and training. Oscar-SW is a novel software-based technique supporting direct AllReduce that splits gradients into probable zeros and non zeros to apply count sketch compression. Oscar-HW, a novel hardware/software codesigned gradient splitting technique is proposed with ASC (Adaptive Stepwise Coding), an encoding technique for gradient compression in distributed training. Oscar-HW dynamically splits fixed-point quantized gradients for AllReduce communications and maximizes bandwidth utilization for state-of-the-art hardware compression techniques. ASC is a variant of Adaptive Arithmetic Coding (AAC) that generates a distinct probability table for each timestep of AllReduce to adapt to its unique value ranges and avoids sending the probability table during the communication of gradients. Our experimental results show that Oscar-SW achieves1.22×1.22 \timesspeedup and 7 % better accuracy over the SOTA CountSketch algorithm. Oscar-HW achieves an average AllReduce speedup of3.77×3.77 \times, and an average end-to-end training speedup of1.38×1.38 \times. ASC achieves an average AllReduce speedup of1.05×1.05 \timesover Atalanta and4.66×4.66 \timesover no compression.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 388b0b75-7eef-47c4-96cd-5d9ada2bdaad

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines