Lune

HPCA2026顶会

Compression-Aware Gradient Splitting for Collective Communications in Distributed Training

Pranati Majhi, Sabuj Laskar, Abdullah Muzahid, Eun Jung Kim

2026年份

摘要

While distributed training is crucial for scaling deep learning models, it incurs significant overhead due to the collective communication of gradients. To alleviate the burden, compression techniques are commonly used to improve network bandwidth utilization. However, compression poses challenges for synchronized AllReduce collective communications, even more so in scalable systems. Non-uniform data sizes resulting from compression can cause bandwidth under-utilization, as faster nodes remain idle while waiting for slower nodes to complete data exchanges, increasing overall communication and consequently, training time. However, the inherent similarity in gradients across consecutive batches presents an opportunity to mitigate these inefficiencies. By leveraging the quantization of gradients and consistent distribution of zeros, the gradients can be partitioned logically to speedup communication. Splitting them into groups with and without zeros can allow different compression approaches for both. The bandwidth under-utilization due to nonuniform data size can also be solved by partitioning the gradients into variable-sized chunks, leading to more balanced compressed data sizes and reduced idle waiting time. We propose two novel strategies in Oscar, where gradient splitting is designed to improve communication and training. Oscar-SW is a novel software-based technique supporting direct AllReduce that splits gradients into probable zeros and non zeros to apply count sketch compression. Oscar-HW, a novel hardware/software codesigned gradient splitting technique is proposed with ASC (Adaptive Stepwise Coding), an encoding technique for gradient compression in distributed training. Oscar-HW dynamically splits fixed-point quantized gradients for AllReduce communications and maximizes bandwidth utilization for state-of-the-art hardware compression techniques. ASC is a variant of Adaptive Arithmetic Coding (AAC) that generates a distinct probability table for each timestep of AllReduce to adapt to its unique value ranges and avoids sending the probability table during the communication of gradients. Our experimental results show that Oscar-SW achieves1.22×1.22 \timesspeedup and 7 % better accuracy over the SOTA CountSketch algorithm. Oscar-HW achieves an average AllReduce speedup of3.77×3.77 \times, and an average end-to-end training speedup of1.38×1.38 \times. ASC achieves an average AllReduce speedup of1.05×1.05 \timesover Atalanta and4.66×4.66 \timesover no compression.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 388b0b75-7eef-47c4-96cd-5d9ada2bdaad

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖