Protean: VM Allocation Service at Scale
Ori Hadary, Luke Marshall, Ishai Menache, Abhisek Pan, Esaias E. Greeff, David Dion, Star Dorminey, Shailesh Joshi, Yang Chen, Mark Russinovich, Thomas Moscibroda
Abstract
We describe the design and implementation of Protean - the Microsoft Azure service responsible for allocating Virtual Machines (VMs) to millions of servers around the globe A single instance of Protean serves an entire availability zone (10-100k machines), facilitating seamless failover and scale-out to customers The design has proven robust, enabling a substantial expansion of VM offerings and features with minimal changes to the core infrastructure In particular, Protean preserves a clear separation between policy and mechanisms From a policy perspective, a flexible rule-based Allocation Agent (AA) allows Protean to efficiently address multiple constraints and performance criteria, and adapt to different conditions On the system side, a multi-layer caching mechanism expedites the allocation process, achieving turnaround times of few milliseconds A slight compromise on allocation quality enables multiple AAs to run concurrently on the same inventory, resulting in increased throughput with negligible conflict rate Our results from both simulations and production demonstrate that Protean achieves high throughput and utilization (85-90% on a key utilization metric), while satisfying user-specific requirements We also demonstrate how Protean is adapted to handle capacity crunch conditions, by zooming in on spikes caused by COVID-19 © 2020 Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020 All rights reserved
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91f9d352-9373-48bd-af95-1eb84cde8434Cited by top-tier papers43
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- Llumnix: Dynamic Scheduling for Large Language Model ServingBiao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao et al.OSDI 2024 · 189 citations
- SONIC: Application-aware Data Passing for Chained Serverless ApplicationsAshraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic et al.USENIX ATC 2021 · 170 citations
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang et al.USENIX ATC 2023 · 115 citations
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun et al.OSDI 2020 · 101 citations
Builds on2
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
- Improving resource utilization by timely fine-grained schedulingTatiana Jin, Zhenkun Cai, Boyang Li, Chengguang Zheng et al.EuroSys 2020 · 27 citations
Related papers
- Kerveros: Efficient and Scalable Cloud Admission ControlSultan Mahmud Sajal, Luke Marshall, Beibin Li, Shandan Zhou et al.OSDI 2023 · 9 citations
- SOL: safe on-node learning in cloud platformsYawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger et al.ASPLOS 2022 · 14 citations
- Correlation-Aware Heuristic Search for Intelligent Virtual Machine Provisioning in Cloud SystemsChuan Luo, Bo Qiao, Wenqian Xing, Xin Chen et al.AAAI 2021 · 19 citations
- SelfTune: Tuning Cluster ManagersAjaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao et al.NSDI 2023 · 30 citations
- Bluebird: High-performance SDN for Bare-metal Cloud ServicesManikandan Arumugam, Deepak Bansal, Navdeep Bhatia, James Boerner et al.NSDI 2022
