Scheduling of Time-Varying Workloads Using Reinforcement Learning
Shanka Subhra Mondal, Nikhil Sheoran, Subrata Mitra
摘要
Resource usage of production workloads running on shared compute clusters often fluctuate significantly across time. While simultaneous spike in the resource usage between two workloads running on the same machine can create performance degradation, unused resources in a machine results in wastage and undesirable operational characteristics for a compute cluster. Prior works did not consider such temporal resource fluctuations or their alignment for scheduling decisions. Due to the variety of time-varying workloads and their complex resource usage characteristics, it is challenging to design well-defined heuristics for scheduling them optimally across different machines in a cluster. In this paper, we propose a Deep Reinforcement Learning (DRL) based approach to exploit various temporal resource usage patterns of timevarying workloads as well as a technique for creating equivalence classes among a large number of production workloads to improve scalability of our method. Validations with real production traces from Google and Alibaba show that our technique can significantly improve metrics for operational excellence (e.g. utilization, fragmentation, resource exhaustion etc.) for a cluster compared to the baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- Accelerating Serverless Computing by Harvesting Idle ResourcesHanfei Yu, Hao Wang, Jian Li, Xu Yuan 等WWW 2022 · 被引用 44 次
- CausIL: Causal Graph for Instance Level Microservice DataSarthak Chakraborty, Shaddy Garg, Shubham Agarwal, Ayush Chauhan 等WWW 2023 · 被引用 20 次
- Conditional Generative Model Based Predicate-Aware Query ApproximationNikhil Sheoran, Subrata Mitra, Vibhor Porwal, Siddharth Ghetia 等AAAI 2022 · 被引用 14 次
- Primo: Practical Learning-Augmented Systems with Interpretable ModelsQinghao Hu, Harsha Nori, Peng Sun, Yonggang Wen 等USENIX ATC 2022 · 被引用 5 次
相关 Paper
- Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement LearningQiyang Ding, Pengfei Zheng, Shreyas Kudari, Shivaram Venkataraman 等SC 2023 · 被引用 5 次
- Graph Assisted Offline-Online Deep Reinforcement Learning for Dynamic Workflow SchedulingYifan Yang, Gang Chen, Hui Ma, Cong Zhang 等ICLR 2025
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen 等SC 2021 · 被引用 136 次
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan 等INFOCOM 2022 · 被引用 27 次
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu 等OSDI 2026
