Building Scalable and Flexible Cluster Managers Using Declarative Programming
Lalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi, Nina Narodytska, Leonid Ryzhyk, Sahan Gamage, Brian Oki, Pranshu Jain, Michael Gasch
摘要
Cluster managers like Kubernetes and OpenStack are notoriously hard to develop, given that they routinely grapple with hard combinatorial optimization problems like load balancing, placement, scheduling, and configuration. Today, cluster manager developers tackle these problems by developing system-specific best effort heuristics, which achieve scalability by significantly sacrificing the cluster manager's decision quality, feature set, and extensibility over time. This is proving untenable, as solutions for cluster management problems are routinely developed from scratch in the industry to solve largely similar problems across different settings.
We propose DCM, a radically different architecture where developers specify the cluster manager's behavior declaratively, using SQL queries over cluster state stored in a relational database. From the SQL specification, the DCM compiler synthesizes a program that, at runtime, can be invoked to compute policy-compliant cluster management decisions given the latest cluster state. Under the covers, the generated program efficiently encodes the cluster state as an optimization problem that can be solved using off-the-shelf solvers, freeing developers from having to design ad-hoc heuristics.
We show that DCM significantly lowers the barrier to building scalable and extensible cluster managers. We validate our claim by powering three production-grade systems with it: a Kubernetes scheduler, a virtual machine management solution, and a distributed transactional datastore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Solving Large-Scale Granular Resource Allocation Problems Efficiently with POPDeepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft 等SOSP 2021 · 被引用 56 次
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su 等NeurIPS 2025 · 被引用 35 次
- DBOS: A DBMS-oriented Operating SystemAthinagoras Skiadopoulos, Qian Li, Peter Kraft, Kostis Kaffes 等VLDB 2022 · 被引用 31 次
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar 等SOSP 2021 · 被引用 19 次
- Acto: Automatic End-to-End Testing for Operation Correctness of Cloud System ManagementJiawei Tyler Gu, Xudong Sun, Wentao Zhang, Yuxuan Jiang 等SOSP 2023 · 被引用 17 次
相关 Paper
- Scaling a Declarative Cluster Manager Architecture with Query Optimization TechniquesKexin Rong, Mihai Budiu, Athinagoras Skiadopoulos, Lalith Suresh 等VLDB 2023 · 被引用 3 次
- Solver-In-The-Loop Cluster Resource Management for Database-as-a-ServiceArnd Christian König, Yi Shan, Karan Newatia, Luke Marshall 等VLDB 2023 · 被引用 9 次
- Kivi: Verification for Cluster ManagementBingzhe Liu, Gangmuk Lim, Ryan Beckett, Philip Brighten GodfreyUSENIX ATC 2024 · 被引用 6 次
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan 等OSDI 2022 · 被引用 44 次
- SelfTune: Tuning Cluster ManagersAjaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao 等NSDI 2023 · 被引用 30 次
