A Case for Task Sampling based Learning for Cluster Job Scheduling
Akshay Jajoo, Y. Charlie Hu, Xiaojun Lin, Nan Deng
摘要
The ability to accurately estimate job runtime properties allows a scheduler to effectively schedule jobs. State-of-the-art online cluster job schedulers use history-based learning, which uses past job execution information to estimate the runtime properties of newly arrived jobs. However, with fast-paced development in cluster technology (in both hardware and software) and changing user inputs, job runtime properties can change over time, which lead to inaccurate predictions. In this paper, we explore the potential and limitation of real-time learning of job runtime properties, by proactively sampling and scheduling a small fraction of the tasks of each job. Such a task-sampling-based approach exploits the similarity among runtime properties of the tasks of the same job and is inherently immune to changing job behavior. Our study focuses on two key questions in comparing task-sampling-based learning (learning in space) and history-based learning (learning in time): (1) Can learning in space be more accurate than learning in time? (2) If so, can delaying scheduling the remaining tasks of a job till the completion of sampled tasks be more than compensated by the improved accuracy and result in improved job performance? Our analytical and experimental analysis of 3 production traces with different skew and job distribution shows that learning in space can be substantially more accurate. Our simulation and testbed evaluation on Azure of the two learning approaches anchored in a generic job scheduler using 3 production cluster job traces shows that despite its online overhead, learning in space reduces the average Job Completion Time (JCT) by 1.28x, 1.56x, and 1.32x compared to the prior-art history-based predictor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine LearningYihong Li, Tianyu Zeng, Xiaoxi Zhang, Jingpu Duan 等INFOCOM 2023 · 被引用 39 次
- Serverless Cold Starts and Where to Find ThemArtjom Joosen, Ahmed Hassan, Martin Asenov, Rajkarn Singh 等EuroSys 2025 · 被引用 30 次
- When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric SchedulingAbdullah Bin Faisal, Noah Martin, Hafiz Mohsin Bashir, Swaminathan Lamelas 等OSDI 2024 · 被引用 6 次
相关 Paper
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 被引用 5 次
- SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement LearningDi Zhang, Dong Dai, Bing XieHPDC 2022 · 被引用 26 次
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo 等OSDI 2023 · 被引用 41 次
- A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningMingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan 等KDD 2023 · 被引用 9 次
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 等NSDI 2025 · 被引用 23 次
