MerKury: Adaptive Resource Allocation to Enhance the Kubernetes Performance for Large-Scale Clusters
Jiayin Luo, Xinkui Zhao, Yuxin Ma, Shengye Pang, Jianwei Yin
Abstract
As a prevalent paradigm of modern web applications, cloud computing has experienced a surge in adoption. The deployment of vast and various workloads encapsulated within containers has become ubiquitous across cloud platforms, imposing substantial demands on the supporting infrastructure. However, Kubernetes (k8s), the de-facto standard for container orchestration, struggles with low scheduling throughput and high latency in large-scale clusters. The primary challenges are identified as excessive loads of read requests and resource contention among co-located components. In response to these challenges, in this paper, we present MerKury, a lightweight framework to enhance the Kubernetes performance for large-scale clusters. It employs a dual strategy: first, it preprocesses specific requests to alleviate unnecessary load, and second, it introduces an adaptive resource allocation algorithm to mitigate resource contention. Evaluations under different scenarios of varying cluster scale have demonstrated that MerKury notably augments cluster scheduling throughput up to 16.4× and reduces request latency by up to 39.3%, outperforming vanilla Kubernetes and baseline resource allocation methods. CCS CONCEPTS • Computer systems organization → Cloud computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 765fd56d-43db-4a71-b953-1ea1185d5a79Builds on12
- StepConf: SLO-Aware Dynamic Resource Configuration for Serverless Function WorkflowsZhaojie Wen, Yishuo Wang, Fangming LiuINFOCOM 2022 · 72 citations
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi et al.NSDI 2024 · 54 citations
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan et al.OSDI 2022 · 44 citations
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo et al.OSDI 2023 · 41 citations
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu et al.EuroSys 2023 · 34 citations
Related papers
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang et al.INFOCOM 2021 · 93 citations
- Oakestra: A Lightweight Hierarchical Orchestration Framework for Edge ComputingGiovanni Bartolomeo, Mehdi Yosofie, Simon Bäurle, Oliver Haluszczynski et al.USENIX ATC 2023 · 51 citations
- KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container CloudTing-An Yeh, Hung-Hsin Chen, Jerry ChouHPDC 2020 · 56 citations
- Pyramid: A Secure, Resource-Efficient, and Pluggable Kubernetes for Multi-TenancyXiang Li, Weijie Liu, Fabing Li, Hongliang Tian et al.EuroSys 2026
- vBOIDs: Taming Chaos via Coarse-Grained Scheduling Abstraction for ContainersKaesi Manakkal, Nathan Daughety, Yu Sun, Marcus Pendleton et al.OSDI 2026
