SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine Learning
Hans Kasan, Dennis Abts, Jungwook Choi, John Kim
Abstract
The interconnection network is a critical component for building scalable systems, as its communication bandwidth directly impacts the collective communication performance of distributed training.In this work, we exploit interconnection network sparsity (or communication sparsity) to address challenges of communication performance and scalability.In particular, we identify how gradients (or packets) during communication can be randomly skipped with minimal impact on accuracy.However, skipping gradients in fine granularity (or individually) results in a loss of gradient information without improving communication performance, due to the synchronous nature of collective communication.Thus, we propose coarse-grained skipping where gradient slices are skipped, which enables skipping of some AllReduce steps to accelerate communication.In particular, we propose SkipReduce collective communication that intentionally skips random gradients during AllReduce.However, a naive implementation of SkipReduce can degrade accuracy by repeatedly skipping gradients from the same node, which introduces bias.To mitigate this accuracy loss, we show how randomizing the skipped gradient slices improves training accuracy with negligible additional runtime.We also observe that not all layers have similar communication sparsity and propose applying SkipReduce selectively where only the sparse layers (or gradients) are skipped to minimize the accuracy impact of SkipReduce.Compared to prior work on communication acceleration, SkipReduce can be seamlessly integrated into existing collective communication libraries with minimal overhead.We implement SkipReduce on top of NCCL's ring-based AllReduce algorithm.Our results show that this method accelerates collective communication while preserving final training accuracy.Compared to baseline AllReduce, SkipReduce provides up to a 1.58× speedup in time-to-accuracy.Beyond this performance gain in data parallelism, this work also * Part of this work was conducted during an internship at NVIDIA Research.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 180e7021-8d7a-4d4d-af1c-5a6f90c4d798Related papers
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- Logical/Physical Topology-Aware Collective Communication in Deep Learning TrainingJo Sanghoon, Hyojun Son, John KimHPCA 2023 · 23 citations
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 26 citations
