AutoSync: Learning to Synchronize for Data-Parallel Distributed Deep Learning
Hao Zhang, Yuan Li, Zhijie Deng, Xiaodan Liang, Lawrence Carin, Eric P. Xing
Abstract
Synchronization is a key step in data-parallel distributed machine learning (ML). Different synchronization systems and strategies perform differently, and to achieve optimal parallel training throughput requires synchronization strategies that adapt to model structures and cluster configurations. Existing synchronization systems often only consider a single or a few synchronization aspects, and the burden of deciding the right synchronization strategy is then placed on the ML practitioners, who may lack the required expertise. In this paper, we develop a model-and resource-dependent representation for synchronization, which unifies multiple synchronization aspects ranging from architecture, message partitioning, placement scheme, to communication topology. Based on this representation, we build an endto-end pipeline, AutoSync, to automatically optimize synchronization strategies given model structures and resource specifications, lowering the bar for dataparallel distributed ML. By learning from low-shot data collected in only 200 trial runs, AutoSync can discover synchronization strategies up to 1.6x better than manually optimized ones. We develop transfer-learning mechanisms to further reduce the auto-optimization cost -the simulators can transfer among similar model architectures, among similar cluster configurations, or both. We also present a dataset that contains nearly 10000 strategy and run-time pairs on a diverse set of models and cluster specifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6010e02f-1b29-4fec-9b68-157b97bfccb5Cited by top-tier papers5
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin et al.OSDI 2022 · 105 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
- Sia: Heterogeneity-aware, goodput-optimized ML-cluster schedulingSuhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao et al.SOSP 2023 · 50 citations
- nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning TrainingZhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang et al.OSDI 2024 · 38 citations
- Only Buffer When You Need To: Reducing On-chip GPU Traffic with Reconfigurable Local Atomic BuffersPreyesh Dalmia, Rohan Mahapatra, Matthew D. SinclairHPCA 2022 · 11 citations
Related papers
- UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic ProgrammingHao Lin, Ke Wu, Jie Li, Jun Li et al.CVPR 2025
- Learning Efficient Parameter Server Synchronization Policies for Distributed SGDRong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian et al.ICLR 2020 · 9 citations
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 48 citations
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
- Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN TrainingWeiyan Wang, Cengguang Zhang, Liu Yang, Kai Chen et al.INFOCOM 2022 · 14 citations
