Unearthing inter-job dependencies for better cluster scheduling
Andrew Chung, Subru Krishnan, Konstantinos Karanasos, Carlo Curino, Gregory R. Ganger
Abstract
Inter-job dependencies pervade shared data analytics infrastructures (so-called "data lakes"), as jobs read output files written by previous jobs, yet are often invisible to current cluster schedulers. Jobs are submitted one-by-one, without indicating dependencies, and the scheduler considers them independently based on priority, fairness, etc. This paper analyzes hidden inter-job dependencies in a 50k+ node analytics cluster at Microsoft, based on job and data provenance logs, finding that nearly 80% of all jobs depend on at least one other job. Yet, even in a business-critical setting, we see jobs that fail because they depend on not-yet-completed jobs, jobs that depend on jobs of lower priority, and other difficulties with hidden inter-job dependencies.
The Wing dependency profiler analyzes job and data provenance logs to find hidden inter-job dependencies, characterizes them, and provides improved guidance to a cluster scheduler. Specifically, for the 68% of jobs (in the analyzed data lake) that exhibit their dependencies in a recurring fashion, Wing predicts the impact of a pending job on subsequent jobs and user downloads, and uses that information to refine valuation of that job by the scheduler. In simulations driven by real job logs, we find that a traditional YARN scheduler that uses Wing-provided valuations in place of user-specified priorities extracts more value (in terms of successful dependent jobs and user downloads) from a heavily-loaded cluster. By relying completely on Wing for guidance, YARN can achieve nearly 100% of value at constrained cluster capacities, almost 2× that achieved by using the user-provided job priorities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1723854-3de4-4549-b627-e0c773f68c37Cited by top-tier papers8
- Generating Complex, Realistic Cloud Workloads using Recurrent Neural NetworksShane Bergsma, Timothy Zeyl, Arik Senderovich, J. Christopher BeckSOSP 2021 · 21 citations
- Unlocking unallocated cloud capacity for long, uninterruptible workloadsAnup Agarwal, Shadi A. Noghabi, Íñigo Goiri, Srinivasan Seshan et al.NSDI 2023 · 9 citations
- K9db: Privacy-Compliant Storage For Web Applications By ConstructionKinan Dak Albab, Ishan Sharma, Justus Adam, Benjamin Kilimnik et al.OSDI 2023 · 7 citations
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 5 citations
- Truthful Online Scheduling of Cloud Workloads under UncertaintyMoshe Babaioff, Ronny Lempel, Brendan Lucier, Ishai Menache et al.WWW 2022 · 4 citations
Related papers
- Moirai: Optimizing Placement of Data and Compute in Hybrid CloudsZiyue Qiu, Hojin Park, Jing Zhao, Yu-Kai Wang et al.SOSP 2025
- SPADE: Signal-Aware DAG Scheduling and Dynamic Provisioning for Data Processing ClustersAdam Lechowicz, Rohan Shenoy, Noman Bashir, Mohammad Hajiesmaili et al.OSDI 2026
- A Community Cache with Complete InformationMania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar et al.FAST 2021 · 2 citations
- Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement LearningQiyang Ding, Pengfei Zheng, Shreyas Kudari, Shivaram Venkataraman et al.SC 2023 · 5 citations
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das et al.VLDB 2021 · 10 citations
