Herring: rethinking the parameter server at scale for the cloud
Indu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai, Yauheni Selivonchyk, Cory Pruce
摘要
Training large deep neural networks is time-consuming and may take days or even weeks to complete. Although parameter-server-based approaches were initially popular in distributed training, scalability issues led the field to move towards all-reduce-based approaches. Recent developments in cloud networking technologies, however, such as the Elastic Fabric Adapter (EFA) and Scalable Reliable Datagram (SRD), motivate a re-thinking of the parameter-server approach to address its fundamental inefficiencies. To this end, we introduce a novel communication library, Herring, which is designed to alleviate the performance bottlenecks in parameter-server-based training. We show that gradient reduction with Herring is twice as fast as all-reduce-based methods. We further demonstrate that training deep learning models like using Herring outperforms all-reduce-based training, achieving 85% scaling efficiency on large clusters with up to 2048 NVIDIA V100 GPUs without accuracy drop.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsSaeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan 等ISCA 2022 · 被引用 40 次
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
- FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network TopologyJinho Lee, Inseok Hwang, Soham Shah, Minsik ChoDAC 2020 · 被引用 21 次
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable DevicesMax Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, Gennady PekhimenkoNeurIPS 2021 · 被引用 59 次
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
