Global Capacity Management With Flux
Marius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall, Po-Yen Chou, Richard Cornew, Adela Kabiljo, Ranjith Kumar S., Maroo Lieuw, Justin Meza, Scott Michelson, Thomas Rohloff
Abstract
Customers of both private and public clouds must wrestle with the problem of regionalization: how should service capacity be apportioned across a large number of geo-distributed datacenter regions? This problem is further complicated by the complex service dependency graphs that arise from microservice architectures, as well as capacity availability and hardware mix that can vary greatly by region.
Historically, regionalization has been solved through a slow-moving and manual process, whereby owners of large services directly negotiate capacity allocation and distribution with the cloud provider. However, as both service and cloud footprints continue to grow, these manual processes are becoming untenable, and often result in excessive labor for all parties involved, as well as suboptimal outcomes.
At Meta, we have built a system called Flux to automate capacity regionalization, transitioning it from a bottoms-up, manual process, to a top-down, automated one. Flux employs RPC tracing to identify service capacity models, and uses these to compute an optimal joint capacity and traffic distribution plan that spans thousands of services across tens of products, and involves millions of servers. These plans are orchestrated by a system that safely and efficiently rebalances service capacity and product traffic across tens of regions on a continuous basis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72d61d28-614e-4c06-9992-a5888e7c8d21Cited by top-tier papers2
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang et al.SOSP 2023 · 32 citations
- Cooperative Graceful Degradation in Containerized CloudsKapil Agrawal, Sangeetha Abdu JyothiASPLOS 2025 · 2 citations
Builds on5
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflowsDarby Huye, Yuri Shkuro, Raja R. SambasivanUSENIX ATC 2023 · 65 citations
- Building Scalable and Flexible Cluster Managers Using Declarative ProgrammingLalith Suresh, João Loff, Faria Kalim, Sangeetha Abdu Jyothi et al.OSDI 2020 · 23 citations
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar et al.SOSP 2021 · 19 citations
- Shard Manager: A Generic Shard Management Framework for Geo-distributed ApplicationsSangmin Lee, Zhenhua Guo, Omer Sunercan, Jun Ying et al.SOSP 2021 · 18 citations
Related papers
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan et al.OSDI 2024 · 39 citations
- Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and ExperiencesNeeraj Kumar, Pol Mauri Ruiz, Vijay Menon, Igor Kabiljo et al.OSDI 2024 · 6 citations
- ServiceRouter: Hyperscale and Minimal Cost Service Mesh at MetaHarshit Saokar, Soteris Demetriou, Nick Magerko, Max Kontorovich et al.OSDI 2023
- Defcon: Preventing Overload with Graceful Feature DegradationJustin Meza, Thote Gowda, Ahmed Eid, Tomiwa Ijaware et al.OSDI 2023 · 20 citations
- Klotski: Efficient and Safe Network Migration of Large Production DatacentersYihao Zhao, Xiaoxiang Zhang, Hang Zhu, Ying Zhang et al.SIGCOMM 2023 · 5 citations
