Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs
Qinghao Hu, Meng Zhang, Peng Sun, Yonggang Wen, Tianwei Zhang
Abstract
While recent deep learning workload schedulers exhibit excellent performance, it is arduous to deploy them in practice due to some substantial defects, including inflexible intrusive manner, exorbitant integration and maintenance cost, limited scalability, as well as opaque decision processes. Motivated by these issues, we design and implement Lucid, a non-intrusive deep learning workload scheduler based on interpretable models. It consists of three innovative modules. First, a two-dimensional optimized profiler is introduced for efficient job metric collection and timely debugging job feedback. Second, Lucid utilizes an indolent packing strategy to circumvent interference. Third, Lucid orchestrates resources based on estimated job priority values and sharing scores to achieve efficient scheduling. Additionally, Lucid promotes model performance maintenance and system transparent adjustment via a well-designed system optimizer. Our evaluation shows that Lucid reduces the average job completion time by up to 1.3× compared with state-of-the-art preemptive scheduler Tiresias. Furthermore, it provides explicit system interpretations and excellent scalability for practical deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3df6f6c-7d40-4a7a-9c87-ba263786a55dCited by top-tier papers14
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- Design and Operation of Shared Machine Learning Clusters on CampusKaiqiang Xu, Decang Sun, Hao Wang, Zhenghang Ren et al.ASPLOS 2025 · 24 citations
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang et al.NSDI 2025 · 23 citations
- vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingJehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim et al.MICRO 2024 · 16 citations
- Hydro: Surrogate-Based Hyperparameter Tuning Service in DatacentersQinghao Hu, Zhisheng Ye, Meng Zhang, Qiaoling Chen et al.OSDI 2023 · 16 citations
Builds on28
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
Related papers
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin et al.NSDI 2021 · 111 citations
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams et al.SC 2020 · 21 citations
- Multi-resource interleaving for deep learning trainingYihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu et al.SIGCOMM 2022 · 78 citations
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
- Blox: A Modular Toolkit for Deep Learning SchedulersSaurabh Agarwal, Amar Phanishayee, Shivaram VenkataramanEuroSys 2024 · 10 citations
