Metis: learning to schedule long-running applications in shared container clusters at scale
Luping Wang, Qizhen Weng, Wei Wang, Chen Chen, Bo Li
摘要
Online cloud services are increasingly deployed as long-running applications (LRAs) in containers. Placing LRA containers is known to be difficult as they often have sophisticated resource interferences and I/O dependencies. Existing schedulers rely on operators to manually express the container scheduling requirements as placement constraints and strive to satisfy as many constraints as possible. Such schedulers, however, fall short in performance as placement constraints only provide qualitative scheduling guidelines and minimizing constraint violations does not necessarily result in the optimal performance.In this work, we present Metis, a general-purpose scheduler that learns to optimally place LRA containers using deep reinforcement learning (RL) techniques. This eliminates the complex manual specification of placement constraints and offers, for the first time, concrete quantitative scheduling criteria. As directly training an RL agent does not scale, we develop a novel hierarchical learning technique that decomposes a complex container placement problem into a hierarchy of subproblems with significantly reduced state and action space. We show that many subproblems have similar structures and can hence be solved by training a unified RL agent offline. Large-scale EC2 deployment shows that compared with the traditional constraint-based schedulers, Metis improves the throughput by up to 61%, optimizes various performance metrics, and easily scales to a large cluster where 3K containers run on over 700 machines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training JobsQinghao Hu, Meng Zhang, Peng Sun, Yonggang Wen 等ASPLOS 2023 · 被引用 45 次
- Accelerating Serverless Computing by Harvesting Idle ResourcesHanfei Yu, Hao Wang, Jian Li, Xu Yuan 等WWW 2022 · 被引用 44 次
- GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsSiling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin 等HPCA 2025 · 被引用 3 次
- MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU ClustersQizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang 等NSDI 2022
相关 Paper
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan 等INFOCOM 2022 · 被引用 27 次
- Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement LearningQiyang Ding, Pengfei Zheng, Shreyas Kudari, Shivaram Venkataraman 等SC 2023 · 被引用 5 次
- A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningMingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan 等KDD 2023 · 被引用 9 次
- Generalizable Resource Allocation in Stream Processing via Deep Reinforcement LearningXiang Ni, Jing Li, Mo Yu, Wang Zhou 等AAAI 2020 · 被引用 24 次
- AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud SystemsHaoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke 等USENIX ATC 2023 · 被引用 95 次
