ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration
Hairui Zhao, Hongliang Li, Qi Tian, Jie Wu, Meng Zhang, Zhewen Xu, Xiang Li, Haixiao Xu
Abstract
Deep Learning (DL) applications have experienced exponential growth in data volume and model complexity, spurring various parallel approaches. Existing solutions mostly focus on accelerating individual training jobs. However, jobs submitted to a cluster may not always be independent. This is due to the distinctive characteristic of DL training that it is an exploratory process. Model developers often launch multiple training instances in a batch with the same model structure but different settings to tune hyper-parameters, which provides an opportunity to regard these jobs as job-arrays. With further support of low-cost job context switching, sharing resources among these jobs is not just feasible but also beneficial to the resource utilization and the throughput of a DL cluster. This paper introduces Job-Array Pipeline Parallelism (JAP) that assembles a batch of sibling DL training jobs into a concurrent job-array. We design ArrayPipe, a framework that supports high throughput model exploration with JAP. A novel scheduling problem in JAP is proposed that seeks to minimize the per-iteration training time for a job-array, along with two scheduling algorithms for different scales of job-arrays. Extensive testbed experiments and trace-driven simulations show that ArrayPipe achieves 1.46× training throughput on average compared with state-of-the-art related works.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6ad2b4bd-48e3-4200-8e7a-dc05de8328b2Cited by top-tier papers4
- FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsHairui Zhao, Qi Tian, Hongliang Li, Zizhong ChenUSENIX ATC 2025 · 6 citations
- TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPUShixun Wu, Yujia Zhai, Huangliang Dai, Yue Zhu et al.SC 2025 · 5 citations
- PASO: Step Parallel Stochastic OptimizationJianrong Lu, Zhuoya Gu, Haobo Li, Zhiyu Zhu et al.ICML 2026 · 1 citation
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun et al.PPoPP 2026
Related papers
- SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel TrainingXupeng Miao, Yining Shi, Zhi Yang, Bin Cui et al.VLDB 2023 · 48 citations
- HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data ParallelismJay H. Park, Gyeongchan Yun, Chang M. Yi, Nguyen T. Nguyen et al.USENIX ATC 2020 · 178 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- Multi-resource interleaving for deep learning trainingYihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu et al.SIGCOMM 2022 · 78 citations
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 10 citations
