FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network Topology
Jinho Lee, Inseok Hwang, Soham Shah, Minsik Cho
Abstract
We propose FlexReduce, an efficient and flexible all-reduce algorithm for distributed deep learning under irregular network hierarchies. With ever-growing deep neural networks, distributed learning over multiple nodes is becoming imperative for expedited training. There are several approaches leveraging the symmetric network structure to optimize the performance over different hierarchy levels of the network. However, the assumption of symmetric network does not always hold, especially in shared cloud environments. By allocating an uneven portion of gradients to each learner (GPU), FlexReduce outperforms conventional algorithms on asymmetric network structures, and still performs even or better on symmetric networks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 92da9a61-2eab-4fbc-b7ae-c22afa95a3edCited by top-tier papers3
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang et al.ASPLOS 2023 · 37 citations
- vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingJehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim et al.MICRO 2024 · 16 citations
- PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesSi Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park et al.ISCA 2024 · 12 citations
Related papers
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
- Logical/Physical Topology-Aware Collective Communication in Deep Learning TrainingJo Sanghoon, Hyojun Son, John KimHPCA 2023 · 23 citations
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang et al.SIGMOD 2021 · 64 citations
