Unearthing inter-job dependencies for better cluster scheduling
Andrew Chung, Subru Krishnan, Konstantinos Karanasos, Carlo Curino, Gregory R. Ganger
摘要
Inter-job dependencies pervade shared data analytics infrastructures (so-called "data lakes"), as jobs read output files written by previous jobs, yet are often invisible to current cluster schedulers. Jobs are submitted one-by-one, without indicating dependencies, and the scheduler considers them independently based on priority, fairness, etc. This paper analyzes hidden inter-job dependencies in a 50k+ node analytics cluster at Microsoft, based on job and data provenance logs, finding that nearly 80% of all jobs depend on at least one other job. Yet, even in a business-critical setting, we see jobs that fail because they depend on not-yet-completed jobs, jobs that depend on jobs of lower priority, and other difficulties with hidden inter-job dependencies.
The Wing dependency profiler analyzes job and data provenance logs to find hidden inter-job dependencies, characterizes them, and provides improved guidance to a cluster scheduler. Specifically, for the 68% of jobs (in the analyzed data lake) that exhibit their dependencies in a recurring fashion, Wing predicts the impact of a pending job on subsequent jobs and user downloads, and uses that information to refine valuation of that job by the scheduler. In simulations driven by real job logs, we find that a traditional YARN scheduler that uses Wing-provided valuations in place of user-specified priorities extracts more value (in terms of successful dependent jobs and user downloads) from a heavily-loaded cluster. By relying completely on Wing for guidance, YARN can achieve nearly 100% of value at constrained cluster capacities, almost 2× that achieved by using the user-provided job priorities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Generating Complex, Realistic Cloud Workloads using Recurrent Neural NetworksShane Bergsma, Timothy Zeyl, Arik Senderovich, J. Christopher BeckSOSP 2021 · 被引用 21 次
- Unlocking unallocated cloud capacity for long, uninterruptible workloadsAnup Agarwal, Shadi A. Noghabi, Íñigo Goiri, Srinivasan Seshan 等NSDI 2023 · 被引用 9 次
- K9db: Privacy-Compliant Storage For Web Applications By ConstructionKinan Dak Albab, Ishan Sharma, Justus Adam, Benjamin Kilimnik 等OSDI 2023 · 被引用 7 次
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 被引用 5 次
- Truthful Online Scheduling of Cloud Workloads under UncertaintyMoshe Babaioff, Ronny Lempel, Brendan Lucier, Ishai Menache 等WWW 2022 · 被引用 4 次
相关 Paper
- Moirai: Optimizing Placement of Data and Compute in Hybrid CloudsZiyue Qiu, Hojin Park, Jing Zhao, Yu-Kai Wang 等SOSP 2025
- SPADE: Signal-Aware DAG Scheduling and Dynamic Provisioning for Data Processing ClustersAdam Lechowicz, Rohan Shenoy, Noman Bashir, Mohammad Hajiesmaili 等OSDI 2026
- A Community Cache with Complete InformationMania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar 等FAST 2021 · 被引用 2 次
- Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement LearningQiyang Ding, Pengfei Zheng, Shreyas Kudari, Shivaram Venkataraman 等SC 2023 · 被引用 5 次
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das 等VLDB 2021 · 被引用 10 次
