ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training
Chia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui, Pin-Yu Chen, Xiao Sun, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Zhang, Kailash Gopalakrishnan
摘要
Large-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms is expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques have been proposed and have demonstrated high compression ratios. However, most existing methods do not scale well to large scale distributed systems (due to gradient build-up) and/or fail to evaluate model fidelity (test accuracy) on large datasets. To mitigate these issues, we propose a new compression technique, Scalable Sparsified Gradient Compression (ScaleCom), that leverages similarity in the gradient distribution amongst learners to provide significantly improved scalability. Using theoretical analysis, we show that ScaleCom provides favorable convergence guarantees and is compatible with gradient all-reduce techniques. Furthermore, we experimentally demonstrate that ScaleCom has small overheads, directly reduces gradient traffic and provides high compression rates (65-400X) and excellent scalability (up to 64 learners and 8-12X larger batch sizes over standard training) across a wide range of applications (image, language, and speech) without significant accuracy loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- DeepReduce: A Sparse-tensor Communication Framework for Federated Deep LearningHang Xu, Kelly Kostopoulou, Aritra Dutta, Xin Li 等NeurIPS 2021 · 被引用 48 次
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and AggregationBoxin Wang, Fan Wu, Yunhui Long, Luka Rimanic 等CCS 2021 · 被引用 45 次
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang 等ASPLOS 2023 · 被引用 37 次
- Fine-tuning Language Models over Slow Networks using Activation Quantization with GuaranteesJue Wang, Binhang Yuan, Luka Rimanic, Yongjun He 等NeurIPS 2022 · 被引用 37 次
- Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemHongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park 等HPCA 2024 · 被引用 26 次
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Quantized Compressive Sampling of Stochastic Gradients for Efficient Communication in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 被引用 32 次
相关 Paper
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He 等ICDE 2023 · 被引用 9 次
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep LearningYisu Wang, Ruilong Wu, Xinjiao Li, Dirk KutscherDAC 2025 · 被引用 3 次
- Communication Efficient SGD via Gradient Sampling With Bayes PriorLiuyihan Song, Kang Zhao, Pan Pan, Yu Liu 等CVPR 2021
- COMPSO: Optimizing Gradient Compression for Distributed Training with Second-Order OptimizersBaixi Sun, Weijin Liu, J. Gregory Pauloski, Jiannan Tian 等PPoPP 2025 · 被引用 8 次
- DAGC: Data-Aware Adaptive Gradient CompressionRongwei Lu, Jiajun Song, Bin Chen, Laizhong Cui 等INFOCOM 2023 · 被引用 12 次
