Twine: A Unified Cluster Management System for Shared Infrastructure
Chunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor, Scott Michelson, Thawan Kooburat, Aravind Anbudurai, Matthew Clark, Kabir Gogia, Long Cheng, Ben Christensen, Alex Gartrell
Abstract
We present Twine, Facebook's cluster management system which has been running in production for the past decade. Twine has helped convert our infrastructure from a collection of siloed pools of customized machines dedicated to individual workloads, into a large-scale shared infrastructure with fungible hardware.
Our goal of ubiquitous shared infrastructure leads us to some decisions counter to common practices. For instance, rather than deploying an isolated control plane per cluster, Twine scales a single control plane to manage one million machines across all data centers in a geographic region and transparently move jobs across clusters.
Twine accommodates workload-specific customization in shared infrastructure, and this approach further departs from common practices. The TaskControl API allows an application to collaborate with Twine to handle container lifecycle events, e.g., restarting a ZooKeeper deployment's followers first and its leader last during a rolling upgrade. Host profiles capture hardware and OS settings that workloads can tune to improve performance and reliability; Twine dynamically allocates machines to workloads and switches host profiles accordingly.
Finally, going against the conventional wisdom of prioritizing stacking workloads on big machines to increase utilization, we universally deploy power-efficient small machines outfit with a single CPU and 64GB RAM to achieve higher performance per watt, and we leverage autoscaling to improve machine utilization.
We describe the design of Twine and share our experience in migrating Facebook's workloads onto shared infrastructure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8edfea19-90cd-466e-8236-551333020f3bCited by top-tier papers42
- Carbon Explorer: A Holistic Framework for Designing Carbon Aware DatacentersBilge Acun, Benjamin C. Lee, Fiodar Kazhamiaka, Kiwan Maeng et al.ASPLOS 2023 · 171 citations
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar et al.ASPLOS 2023 · 74 citations
- Lyra: Elastic Scheduling for Deep Learning ClustersJiamin Li, Hong Xu, Yibo Zhu, Zherui Liu et al.EuroSys 2023 · 59 citations
- Anvil: Verifying Liveness of Cluster Management ControllersXudong Sun, Wenjie Ma, Jiawei Tyler Gu, Zicheng Ma et al.OSDI 2024 · 50 citations
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan et al.OSDI 2022 · 44 citations
Builds on4
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- Virtual Consensus in DelosMahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi et al.OSDI 2020 · 42 citations
- Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at ScaleShaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis et al.OSDI 2020 · 42 citations
Related papers
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar et al.SOSP 2021 · 19 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Scaling Large Production Clusters with Partitioned SynchronizationYihui Feng, Zhi Liu, Yunjian Zhao, Tatiana Jin et al.USENIX ATC 2021 · 23 citations
- Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared ClustersShutian Luo, Jianxiong Liao, Chenyu Lin, Huanle Xu et al.ASPLOS 2025 · 2 citations
- Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User WebsiteUsama Naseer, Luca Niccolini, Udip Pant, Alan Frindell et al.SIGCOMM 2020 · 19 citations
