Lune

INFOCOM2025顶会

ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration

Hairui Zhao, Hongliang Li, Qi Tian, Jie Wu, Meng Zhang, Zhewen Xu, Xiang Li, Haixiao Xu

2025年份
3被引次数
4顶会引用

摘要

Deep Learning (DL) applications have experienced exponential growth in data volume and model complexity, spurring various parallel approaches. Existing solutions mostly focus on accelerating individual training jobs. However, jobs submitted to a cluster may not always be independent. This is due to the distinctive characteristic of DL training that it is an exploratory process. Model developers often launch multiple training instances in a batch with the same model structure but different settings to tune hyper-parameters, which provides an opportunity to regard these jobs as job-arrays. With further support of low-cost job context switching, sharing resources among these jobs is not just feasible but also beneficial to the resource utilization and the throughput of a DL cluster. This paper introduces Job-Array Pipeline Parallelism (JAP) that assembles a batch of sibling DL training jobs into a concurrent job-array. We design ArrayPipe, a framework that supports high throughput model exploration with JAP. A novel scheduling problem in JAP is proposed that seeks to minimize the per-iteration training time for a job-array, along with two scheduling algorithms for different scales of job-arrays. Extensive testbed experiments and trace-driven simulations show that ArrayPipe achieves 1.46× training throughput on average compared with state-of-the-art related works.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖