Protean: VM Allocation Service at Scale
Ori Hadary, Luke Marshall, Ishai Menache, Abhisek Pan, Esaias E. Greeff, David Dion, Star Dorminey, Shailesh Joshi, Yang Chen, Mark Russinovich, Thomas Moscibroda
摘要
We describe the design and implementation of Protean - the Microsoft Azure service responsible for allocating Virtual Machines (VMs) to millions of servers around the globe A single instance of Protean serves an entire availability zone (10-100k machines), facilitating seamless failover and scale-out to customers The design has proven robust, enabling a substantial expansion of VM offerings and features with minimal changes to the core infrastructure In particular, Protean preserves a clear separation between policy and mechanisms From a policy perspective, a flexible rule-based Allocation Agent (AA) allows Protean to efficiently address multiple constraints and performance criteria, and adapt to different conditions On the system side, a multi-layer caching mechanism expedites the allocation process, achieving turnaround times of few milliseconds A slight compromise on allocation quality enables multiple AAs to run concurrently on the same inventory, resulting in increased throughput with negligible conflict rate Our results from both simulations and production demonstrate that Protean achieves high throughput and utilization (85-90% on a key utilization metric), while satisfying user-specific requirements We also demonstrate how Protean is adapted to handle capacity crunch conditions, by zooming in on spikes caused by COVID-19 © 2020 Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020 All rights reserved
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- Llumnix: Dynamic Scheduling for Large Language Model ServingBiao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao 等OSDI 2024 · 被引用 189 次
- SONIC: Application-aware Data Passing for Chained Serverless ApplicationsAshraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic 等USENIX ATC 2021 · 被引用 170 次
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang 等USENIX ATC 2023 · 被引用 115 次
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun 等OSDI 2020 · 被引用 101 次
它引用的顶会 Paper2
相关 Paper
- Kerveros: Efficient and Scalable Cloud Admission ControlSultan Mahmud Sajal, Luke Marshall, Beibin Li, Shandan Zhou 等OSDI 2023 · 被引用 9 次
- SOL: safe on-node learning in cloud platformsYawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger 等ASPLOS 2022 · 被引用 14 次
- Correlation-Aware Heuristic Search for Intelligent Virtual Machine Provisioning in Cloud SystemsChuan Luo, Bo Qiao, Wenqian Xing, Xin Chen 等AAAI 2021 · 被引用 19 次
- SelfTune: Tuning Cluster ManagersAjaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao 等NSDI 2023 · 被引用 30 次
- Bluebird: High-performance SDN for Bare-metal Cloud ServicesManikandan Arumugam, Deepak Bansal, Navdeep Bhatia, James Boerner 等NSDI 2022
