Scheduling of Time-Varying Workloads Using Reinforcement Learning
Shanka Subhra Mondal, Nikhil Sheoran, Subrata Mitra
Abstract
Resource usage of production workloads running on shared compute clusters often fluctuate significantly across time. While simultaneous spike in the resource usage between two workloads running on the same machine can create performance degradation, unused resources in a machine results in wastage and undesirable operational characteristics for a compute cluster. Prior works did not consider such temporal resource fluctuations or their alignment for scheduling decisions. Due to the variety of time-varying workloads and their complex resource usage characteristics, it is challenging to design well-defined heuristics for scheduling them optimally across different machines in a cluster. In this paper, we propose a Deep Reinforcement Learning (DRL) based approach to exploit various temporal resource usage patterns of timevarying workloads as well as a technique for creating equivalence classes among a large number of production workloads to improve scalability of our method. Validations with real production traces from Google and Alibaba show that our technique can significantly improve metrics for operational excellence (e.g. utilization, fragmentation, resource exhaustion etc.) for a cluster compared to the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb4fc034-6761-4b0a-8d7e-c81585de3461Cited by top-tier papers6
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Accelerating Serverless Computing by Harvesting Idle ResourcesHanfei Yu, Hao Wang, Jian Li, Xu Yuan et al.WWW 2022 · 44 citations
- CausIL: Causal Graph for Instance Level Microservice DataSarthak Chakraborty, Shaddy Garg, Shubham Agarwal, Ayush Chauhan et al.WWW 2023 · 20 citations
- Conditional Generative Model Based Predicate-Aware Query ApproximationNikhil Sheoran, Subrata Mitra, Vibhor Porwal, Siddharth Ghetia et al.AAAI 2022 · 14 citations
- Primo: Practical Learning-Augmented Systems with Interpretable ModelsQinghao Hu, Harsha Nori, Peng Sun, Yonggang Wen et al.USENIX ATC 2022 · 5 citations
Related papers
- Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement LearningQiyang Ding, Pengfei Zheng, Shreyas Kudari, Shivaram Venkataraman et al.SC 2023 · 5 citations
- Graph Assisted Offline-Online Deep Reinforcement Learning for Dynamic Workflow SchedulingYifan Yang, Gang Chen, Hui Ma, Cong Zhang et al.ICLR 2025
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan et al.INFOCOM 2022 · 27 citations
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu et al.OSDI 2026
