RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation
Andrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar, Maxim Khutornenko, Mayank Pundir, Yirui Zhang, Mingjun Zhang, Yuanlai Liu, Linh Le, Brendon Daugherty, Apurva Samudra
摘要
Capacity reservation is a common offering in public clouds and on-premise infrastructure. However, no prior work provides capacity reservation with SLO guarantees that takes into account random and correlated hardware failures, datacenter maintenance, and heterogeneous hardware. In this paper, we describe how Facebook's region-scale Resource Allowance System (RAS) addresses these issues and provides guaranteed capacity. RAS uses a capacity abstraction called reservation to represent a set of servers dynamically assigned to a logical cluster. We take a two-level approach to scale resource allocation to all datacenters in a region, where a mixed-integer-programming solver continuously optimizes server-to-reservation assignments off the critical path, and a traditional container allocator does real-time placement of containers on servers in a reservation. As a relatively new component of Facebook's 10-year old cluster manager Twine, RAS has been running in production for almost two years, continuously optimizing the allocation of millions of servers to thousands of reservations. We describe the design of RAS and share our experience of deploying it at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNsJohn Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao 等NSDI 2023 · 被引用 144 次
- Solving Large-Scale Granular Resource Allocation Problems Efficiently with POPDeepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft 等SOSP 2021 · 被引用 56 次
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi 等NSDI 2024 · 被引用 54 次
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan 等OSDI 2024 · 被引用 39 次
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo 等OSDI 2022 · 被引用 31 次
它引用的顶会 Paper4
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun 等OSDI 2020 · 被引用 101 次
- Building Scalable and Flexible Cluster Managers Using Declarative ProgrammingLalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi 等OSDI 2020 · 被引用 23 次
相关 Paper
- Solver-In-The-Loop Cluster Resource Management for Database-as-a-ServiceArnd Christian König, Yi Shan, Karan Newatia, Luke Marshall 等VLDB 2023 · 被引用 9 次
- Network entitlement: contract-based network sharing with agility and SLO guaranteesSatyajeet Singh Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram 等SIGCOMM 2022 · 被引用 3 次
- LaSS: Running Latency Sensitive Serverless Computations at the EdgeBin Wang, Ahmed Ali-Eldin, Prashant J. ShenoyHPDC 2021 · 被引用 74 次
- Global Capacity Management With FluxMarius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall 等OSDI 2023 · 被引用 9 次
- Eva: Cost-Efficient Cloud-Based Cluster SchedulingTzu-Tao Chang, Shivaram VenkataramanEuroSys 2025 · 被引用 2 次
