Improving resource utilization by timely fine-grained scheduling
Tatiana Jin, Zhenkun Cai, Boyang Li, Chengguang Zheng, Guanxian Jiang, James Cheng
Abstract
Monotask is a unit of work that uses only a single type of resource (e.g., CPU, network, disk I/O). While monotask was primarily introduced as a means to reason about job performance, in this paper we show that this fine-grained, resource-oriented abstraction can be leveraged by job schedulers to maximize cluster resource utilization. Although recent cluster schedulers have significantly improved resource allocation, the utilization of the allocated resources is often not high due to inaccurate resource requests. In particular, we show that existing scheduling mechanisms are ineffective for handling jobs with dynamic resource usage, which exists in common workloads, and propose a resource negotiation mechanism between job schedulers and executors that makes use of monotasks. We design a new framework, called Ursa, which enables the scheduler to capture accurate resource demands dynamically from the execution runtime and to provide timely, fine-grained resource allocation based on monotasks. Ursa also enables high utilization of the allocated resources by the execution runtime. We show by experiments that Ursa is able to improve cluster resource utilization, which effectively translates to improved makespan and average JCT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song et al.VLDB 2022 · 107 citations
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin et al.EuroSys 2021 · 60 citations
- Seastar: vertex-centric programming for graph neural networksYidi Wu, Kaihao Ma, Zhenkun Cai, Tatiana Jin et al.EuroSys 2021 · 57 citations
- SelfTune: Tuning Cluster ManagersAjaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao et al.NSDI 2023 · 30 citations
Related papers
- Scaling Large Production Clusters with Partitioned SynchronizationYihui Feng, Zhi Liu, Yunjian Zhao, Tatiana Jin et al.USENIX ATC 2021 · 23 citations
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair et al.SC 2024 · 10 citations
- Starburst: A Cost-aware Scheduler for Hybrid CloudMichael Luo, Siyuan Zhuang, Suryaprakash Vengadesan, Romil Bhardwaj et al.USENIX ATC 2024 · 10 citations
- Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionEdward Guo, Weiwei Jia, Xiaoning Ding, Jianchen ShanEuroSys 2025 · 3 citations
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo et al.OSDI 2023 · 41 citations
