Lune

MICRO2025顶会

SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine Learning

Hans Kasan, Dennis Abts, Jungwook Choi, John Kim

2025年份
1被引次数

摘要

The interconnection network is a critical component for building scalable systems, as its communication bandwidth directly impacts the collective communication performance of distributed training.In this work, we exploit interconnection network sparsity (or communication sparsity) to address challenges of communication performance and scalability.In particular, we identify how gradients (or packets) during communication can be randomly skipped with minimal impact on accuracy.However, skipping gradients in fine granularity (or individually) results in a loss of gradient information without improving communication performance, due to the synchronous nature of collective communication.Thus, we propose coarse-grained skipping where gradient slices are skipped, which enables skipping of some AllReduce steps to accelerate communication.In particular, we propose SkipReduce collective communication that intentionally skips random gradients during AllReduce.However, a naive implementation of SkipReduce can degrade accuracy by repeatedly skipping gradients from the same node, which introduces bias.To mitigate this accuracy loss, we show how randomizing the skipped gradient slices improves training accuracy with negligible additional runtime.We also observe that not all layers have similar communication sparsity and propose applying SkipReduce selectively where only the sparse layers (or gradients) are skipped to minimize the accuracy impact of SkipReduce.Compared to prior work on communication acceleration, SkipReduce can be seamlessly integrated into existing collective communication libraries with minimal overhead.We implement SkipReduce on top of NCCL's ring-based AllReduce algorithm.Our results show that this method accelerates collective communication while preserving final training accuracy.Compared to baseline AllReduce, SkipReduce provides up to a 1.58× speedup in time-to-accuracy.Beyond this performance gain in data parallelism, this work also * Part of this work was conducted during an internship at NVIDIA Research.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖