Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences
Neeraj Kumar, Pol Mauri Ruiz, Vijay Menon, Igor Kabiljo, Mayank Pundir, Andrew Newell, Daniel Lee, Liyuan Wang, Chunqiang Tang
摘要
Meta's private cloud uses millions of servers to host tens of thousands of services that power multiple products for billions of users. This complex environment has various optimization problems involving resource allocation, including hardware placement, server allocation, ML training & inference placement, traffic routing, database & container migration for load balancing, grouping serverless functions for locality, etc.
The main challenges for a reusable resource-allocation framework are its usability and scalability. Usability is impeded by practitioners struggling to translate real-life policies into precise mathematical formulas required by formal optimization methods, while scalability is hampered by NPhard problems that cannot be solved efficiently by commercial solvers.
These challenges are addressed by Rebalancer, Meta's resource-allocation framework. It has been applied to dozens of large-scale use cases over the past seven years, demonstrating its usability, scalability, and generality. At the core of Rebalancer is an expression graph that enables its optimization algorithm to run more efficiently than past algorithms. Moreover, Rebalancer offers a high-level specification language to lower the barrier for adoption by systems practitioners.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Oasis: Pooling PCIe Devices Over CXL to Boost UtilizationYuhong Zhong, Daniel S. Berger, Pantea Zardoshti, Enrique Saurez 等SOSP 2025 · 被引用 2 次
- Decisionhouse: Prescriptive Analytics in the Data StackMatteo Brucato, Fjodor Kholodkov, Soren Little, Jakob Mayer 等VLDB 2026 · 被引用 2 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu 等EuroSys 2026
- COpter: Efficient Large-Scale Resource-Allocation via Continual OptimizationSuhas Jayaram Subramanya, Don Kurian Dennis, Virginia Smith, Gregory R. GangerSOSP 2025
它引用的顶会 Paper8
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Solving Large-Scale Granular Resource Allocation Problems Efficiently with POPDeepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft 等SOSP 2021 · 被引用 56 次
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang 等SOSP 2023 · 被引用 32 次
- Building Scalable and Flexible Cluster Managers Using Declarative ProgrammingLalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi 等OSDI 2020 · 被引用 23 次
相关 Paper
- Global Capacity Management With FluxMarius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall 等OSDI 2023 · 被引用 9 次
- Decouple and Decompose: Scaling Resource Allocation with DeDeZhiying Xu, Minlan Yu, Francis Y. YanOSDI 2025 · 被引用 5 次
- ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production TestingMike Chow, Yang Wang, William Wang, Ayichew Hailu 等OSDI 2024 · 被引用 11 次
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan 等OSDI 2024 · 被引用 39 次
- Pocket: ML Serving from the EdgeMisun Park, Ketan Bhardwaj, Ada GavrilovskaEuroSys 2023 · 被引用 10 次
