Global Capacity Management With Flux
Marius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall, Po-Yen Chou, Richard Cornew, Adela Kabiljo, Ranjith Kumar S., Maroo Lieuw, Justin Meza, Scott Michelson, Thomas Rohloff
摘要
Customers of both private and public clouds must wrestle with the problem of regionalization: how should service capacity be apportioned across a large number of geo-distributed datacenter regions? This problem is further complicated by the complex service dependency graphs that arise from microservice architectures, as well as capacity availability and hardware mix that can vary greatly by region.
Historically, regionalization has been solved through a slow-moving and manual process, whereby owners of large services directly negotiate capacity allocation and distribution with the cloud provider. However, as both service and cloud footprints continue to grow, these manual processes are becoming untenable, and often result in excessive labor for all parties involved, as well as suboptimal outcomes.
At Meta, we have built a system called Flux to automate capacity regionalization, transitioning it from a bottoms-up, manual process, to a top-down, automated one. Flux employs RPC tracing to identify service capacity models, and uses these to compute an optimal joint capacity and traffic distribution plan that spans thousands of services across tens of products, and involves millions of servers. These plans are orchestrated by a system that safely and efficiently rebalances service capacity and product traffic across tens of regions on a continuous basis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang 等SOSP 2023 · 被引用 32 次
- Cooperative Graceful Degradation in Containerized CloudsKapil Agrawal, Sangeetha Abdu JyothiASPLOS 2025 · 被引用 2 次
它引用的顶会 Paper5
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflowsDarby Huye, Yuri Shkuro, Raja R. SambasivanUSENIX ATC 2023 · 被引用 65 次
- Building Scalable and Flexible Cluster Managers Using Declarative ProgrammingLalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi 等OSDI 2020 · 被引用 23 次
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar 等SOSP 2021 · 被引用 19 次
- Shard Manager: A Generic Shard Management Framework for Geo-distributed ApplicationsSangmin Lee, Zhenhua Guo, Omer Sunercan, Jun Ying 等SOSP 2021 · 被引用 18 次
相关 Paper
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan 等OSDI 2024 · 被引用 39 次
- Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and ExperiencesNeeraj Kumar, Pol Mauri Ruiz, Vijay Menon, Igor Kabiljo 等OSDI 2024 · 被引用 6 次
- ServiceRouter: Hyperscale and Minimal Cost Service Mesh at MetaHarshit Saokar, Soteris Demetriou, Nick Magerko, Max Kontorovich 等OSDI 2023
- Defcon: Preventing Overload with Graceful Feature DegradationJustin Meza, Thote Gowda, Ahmed Eid, Tomiwa Ijaware 等OSDI 2023 · 被引用 20 次
- Klotski: Efficient and Safe Network Migration of Large Production DatacentersYihao Zhao, Xiaoxiang Zhang, Hang Zhu, Ying Zhang 等SIGCOMM 2023 · 被引用 5 次
