Piper: Multidimensional Planner for DNN Parallelization
Jakub Tarnawski, Deepak Narayanan, Amar Phanishayee
Abstract
The rapid increase in sizes of state-of-the-art DNN models, and consequently the increase in the compute and memory requirements of model training, has led to the development of many execution schemes such as data parallelism, pipeline model parallelism, tensor (intra-layer) model parallelism, and various memory-saving optimizations. However, no prior work has tackled the highly complex problem of optimally partitioning the DNN computation graph across many accelerators while combining all these parallelism modes and optimizations. In this work, we introduce Piper, an efficient optimization algorithm for this problem that is based on a two-level dynamic programming approach. Our two-level approach is driven by the insight that being given tensor-parallelization techniques for individual layers (e.g., Megatron-LM's splits for transformer layers) significantly reduces the search space and makes the global problem tractable, compared to considering tensor-parallel configurations for the entire DNN operator graph. Combining these dimensions, however, is non-trivial [19] , since each dimension has trade-offs with respect to computational efficiency, amount of communication, and memory footprint. Given the importance of efficient model-parallel training (and inference [7, 4] ), partitioning a model across 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang et al.NeurIPS 2022 · 157 citations
- SpotServe: Serving Generative Large Language Models on Preemptible InstancesXupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi et al.ASPLOS 2024 · 71 citations
- AMP: Automatically Finding Model Parallel Strategies with Heterogeneity AwarenessDacheng Li, Hongyi Wang, Eric P. Xing, Hao ZhangNeurIPS 2022 · 57 citations
- Sia: Heterogeneity-aware, goodput-optimized ML-cluster schedulingSuhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao et al.SOSP 2023 · 50 citations
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang et al.ASPLOS 2023 · 37 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
- Reinforced Genetic Algorithm Learning for Optimizing Computation GraphsAditya Paliwal, Felix Gimeno, Vinod Nair, Yujia Li et al.ICLR 2020 · 70 citations
- Transferable Graph Optimizers for ML CompilersYanqi Zhou, Sudip Roy, AmirAli Abdolrashidi, Daniel Wong et al.NeurIPS 2020 · 63 citations
Related papers
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang et al.DAC 2021 · 13 citations
- Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationGuodong Liu, Youshan Miao, Zhiqi Lin, Xiaoxiang Shi et al.EuroSys 2024 · 16 citations
- MeshSlice: Efficient 2D Tensor Parallelism for Distributed DNN TrainingHyoungwook Nam, Gerasimos Gerogiannis, Josep TorrellasISCA 2025 · 3 citations
- AccPar: Tensor Partitioning for Heterogeneous Deep Learning AcceleratorsLinghao Song, Fan Chen, Youwei Zhuo, Xuehai Qian et al.HPCA 2020 · 62 citations
