Improving resource utilization by timely fine-grained scheduling
Tatiana Jin, Zhenkun Cai, Boyang Li, Chengguang Zheng, Guanxian Jiang, James Cheng
摘要
Monotask is a unit of work that uses only a single type of resource (e.g., CPU, network, disk I/O). While monotask was primarily introduced as a means to reason about job performance, in this paper we show that this fine-grained, resource-oriented abstraction can be leveraged by job schedulers to maximize cluster resource utilization. Although recent cluster schedulers have significantly improved resource allocation, the utilization of the allocated resources is often not high due to inaccurate resource requests. In particular, we show that existing scheduling mechanisms are ineffective for handling jobs with dynamic resource usage, which exists in common workloads, and propose a resource negotiation mechanism between job schedulers and executors that makes use of monotasks. We design a new framework, called Ursa, which enables the scheduler to capture accurate resource demands dynamically from the execution runtime and to provide timely, fine-grained resource allocation based on monotasks. Ursa also enables high utilization of the allocated resources by the execution runtime. We show by experiments that Ursa is able to improve cluster resource utilization, which effectively translates to improved makespan and average JCT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song 等VLDB 2022 · 被引用 107 次
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin 等EuroSys 2021 · 被引用 60 次
- Seastar: vertex-centric programming for graph neural networksYidi Wu, Kaihao Ma, Zhenkun Cai, Tatiana Jin 等EuroSys 2021 · 被引用 57 次
- SelfTune: Tuning Cluster ManagersAjaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao 等NSDI 2023 · 被引用 30 次
相关 Paper
- Scaling Large Production Clusters with Partitioned SynchronizationYihui Feng, Zhi Liu, Yunjian Zhao, Tatiana Jin 等USENIX ATC 2021 · 被引用 23 次
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair 等SC 2024 · 被引用 10 次
- Starburst: A Cost-aware Scheduler for Hybrid CloudMichael Luo, Siyuan Zhuang, Suryaprakash Vengadesan, Romil Bhardwaj 等USENIX ATC 2024 · 被引用 10 次
- Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionEdward Guo, Weiwei Jia, Xiaoning Ding, Jianchen ShanEuroSys 2025 · 被引用 3 次
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo 等OSDI 2023 · 被引用 41 次
