Near-optimal sparse allreduce for distributed deep learning
Shigang Li, Torsten Hoefler
摘要
Communication overhead is one of the major obstacles to train large deep learning models at scale. Gradient sparsification is a promising technique to reduce the communication volume. However, it is very challenging to obtain real performance improvement because of (1) the difficulty of achieving an scalable and efficient sparse allreduce algorithm and (2) the sparsification overhead. This paper proposes Ok-Topk, a scheme for distributed training with sparse gradients. Ok-Topk integrates a novel sparse allreduce algorithm (less than 6k communication volume which is asymptotically optimal) with the decentralized parallel Stochastic Gradient Descent (SGD) optimizer, and its convergence is proved. To reduce the sparsification overhead, Ok-Topk efficiently selects the top-k gradient values according to an estimated threshold. Evaluations are conducted on the Piz Daint supercomputer with neural network models from different deep learning domains. Empirical results show that Ok-Topk achieves similar model accuracy to dense allreduce. Compared with the optimized dense and the state-of-the-art sparse allreduces, Ok-Topk is more scalable and significantly improves training throughput (e.g., 3.29x-12.95x improvement for BERT on 256 GPUs).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang 等ASPLOS 2023 · 被引用 37 次
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 被引用 22 次
- HammingMesh: A Network Topology for Large-Scale Deep LearningTorsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo 等SC 2022 · 被引用 20 次
- CO2: Efficient Distributed Training with Full Communication-Computation OverlapWeigao Sun, Zhen Qin, Weixuan Sun, Shidi Li 等ICLR 2024 · 被引用 17 次
- Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational InequalitiesAleksandr Beznosikov, Martin Takác, Alexander V. GasnikovNeurIPS 2023 · 被引用 15 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen 等ICML 2021 · 被引用 283 次
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao 等PPoPP 2021 · 被引用 224 次
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 被引用 124 次
相关 Paper
- SparDL: Distributed Deep Learning Training with Efficient Sparse CommunicationMinjun Zhao, Yichen Yin, Yuren Mao, Qing Liu 等ICDE 2024 · 被引用 6 次
- SwitchTop-k: Scaling Top-k Compression on Programmable SwitchesYijun Li, Jiawei Huang, Jingling Liu, Zhaoyi Li 等KDD 2025
- ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN TrainingZhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng 等HPDC 2024 · 被引用 5 次
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li 等INFOCOM 2020 · 被引用 66 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
