Efficient sparse collective communication and its application to accelerate distributed deep learning
Jiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini, Amedeo Sapio
摘要
Efficient collective communication is crucial to parallel-computing applications such as distributed training of large-scale recommendation systems and natural language processing models. Existing collective communication libraries focus on optimizing operations for dense inputs, resulting in transmissions of many zeros when inputs are sparse. This counters current trends that see increasing data sparsity in large models.
We propose OmniReduce, an efficient streaming aggregation system that exploits sparsity to maximize effective bandwidth use by sending only non-zero data blocks. We demonstrate that this idea is beneficial and accelerates distributed training by up to 8.2×. Even at 100 Gbps, OmniReduce delivers 1.4-2.9× better performance for network-bottlenecked DNNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated LearningShay Vargaftik, Ran Ben Basat, Amit Portnoy, Gal Mendelson 等ICML 2022 · 被引用 64 次
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 被引用 57 次
- Towards Domain-Specific Network Transport for Distributed DNN TrainingHao Wang, Han Tian, Jingrong Chen, Xinchen Wan 等NSDI 2024 · 被引用 54 次
- Flare: flexible in-network allreduceDaniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li 等SC 2021 · 被引用 49 次
- DeepReduce: A Sparse-tensor Communication Framework for Federated Deep LearningHang Xu, Kelly Kostopoulou, Aritra Dutta, Xin Li 等NeurIPS 2021 · 被引用 48 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
相关 Paper
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 被引用 1 次
- DeMo: Decoupled Momentum OptimizationBowen Peng, Lizhang Chen, Baiyu Su, Jeffrey Quesnelle 等ICLR 2026 · 被引用 7 次
- Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsHaidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai 等SIGIR 2020 · 被引用 9 次
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep LearningYisu Wang, Ruilong Wu, Xinjiao Li, Dirk KutscherDAC 2025 · 被引用 3 次
- Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsSaeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan 等ISCA 2022 · 被引用 40 次
