RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation
Andrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar, Maxim Khutornenko, Mayank Pundir, Yirui Zhang, Mingjun Zhang, Yuanlai Liu, Linh Le, Brendon Daugherty, Apurva Samudra
Abstract
Capacity reservation is a common offering in public clouds and on-premise infrastructure. However, no prior work provides capacity reservation with SLO guarantees that takes into account random and correlated hardware failures, datacenter maintenance, and heterogeneous hardware. In this paper, we describe how Facebook's region-scale Resource Allowance System (RAS) addresses these issues and provides guaranteed capacity. RAS uses a capacity abstraction called reservation to represent a set of servers dynamically assigned to a logical cluster. We take a two-level approach to scale resource allocation to all datacenters in a region, where a mixed-integer-programming solver continuously optimizes server-to-reservation assignments off the critical path, and a traditional container allocator does real-time placement of containers on servers in a reservation. As a relatively new component of Facebook's 10-year old cluster manager Twine, RAS has been running in production for almost two years, continuously optimizing the allocation of millions of servers to thousands of reservations. We describe the design of RAS and share our experience of deploying it at scale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bdecd48-a53d-4079-8411-a0147c2a64a4Cited by top-tier papers18
- Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNsJohn Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao et al.NSDI 2023 · 144 citations
- Solving Large-Scale Granular Resource Allocation Problems Efficiently with POPDeepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft et al.SOSP 2021 · 56 citations
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi et al.NSDI 2024 · 54 citations
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan et al.OSDI 2024 · 39 citations
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo et al.OSDI 2022 · 31 citations
Builds on4
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun et al.OSDI 2020 · 101 citations
- Building Scalable and Flexible Cluster Managers Using Declarative ProgrammingLalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi et al.OSDI 2020 · 23 citations
Related papers
- Solver-In-The-Loop Cluster Resource Management for Database-as-a-ServiceArnd Christian König, Yi Shan, Karan Newatia, Luke Marshall et al.VLDB 2023 · 9 citations
- Network entitlement: contract-based network sharing with agility and SLO guaranteesSatyajeet Singh Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram et al.SIGCOMM 2022 · 3 citations
- LaSS: Running Latency Sensitive Serverless Computations at the EdgeBin Wang, Ahmed Ali-Eldin, Prashant J. ShenoyHPDC 2021 · 74 citations
- Global Capacity Management With FluxMarius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall et al.OSDI 2023 · 9 citations
- Eva: Cost-Efficient Cloud-Based Cluster SchedulingTzu-Tao Chang, Shivaram VenkataramanEuroSys 2025 · 2 citations
