Twine: A Unified Cluster Management System for Shared Infrastructure
Chunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor, Scott Michelson, Thawan Kooburat, Aravind Anbudurai, Matthew Clark, Kabir Gogia, Long Cheng, Ben Christensen, Alex Gartrell
摘要
We present Twine, Facebook's cluster management system which has been running in production for the past decade. Twine has helped convert our infrastructure from a collection of siloed pools of customized machines dedicated to individual workloads, into a large-scale shared infrastructure with fungible hardware.
Our goal of ubiquitous shared infrastructure leads us to some decisions counter to common practices. For instance, rather than deploying an isolated control plane per cluster, Twine scales a single control plane to manage one million machines across all data centers in a geographic region and transparently move jobs across clusters.
Twine accommodates workload-specific customization in shared infrastructure, and this approach further departs from common practices. The TaskControl API allows an application to collaborate with Twine to handle container lifecycle events, e.g., restarting a ZooKeeper deployment's followers first and its leader last during a rolling upgrade. Host profiles capture hardware and OS settings that workloads can tune to improve performance and reliability; Twine dynamically allocates machines to workloads and switches host profiles accordingly.
Finally, going against the conventional wisdom of prioritizing stacking workloads on big machines to increase utilization, we universally deploy power-efficient small machines outfit with a single CPU and 64GB RAM to achieve higher performance per watt, and we leverage autoscaling to improve machine utilization.
We describe the design of Twine and share our experience in migrating Facebook's workloads onto shared infrastructure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Carbon Explorer: A Holistic Framework for Designing Carbon Aware DatacentersBilge Acun, Benjamin C. Lee, Fiodar Kazhamiaka, Kiwan Maeng 等ASPLOS 2023 · 被引用 171 次
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar 等ASPLOS 2023 · 被引用 74 次
- Lyra: Elastic Scheduling for Deep Learning ClustersJiamin Li, Hong Xu, Yibo Zhu, Zherui Liu 等EuroSys 2023 · 被引用 59 次
- Anvil: Verifying Liveness of Cluster Management ControllersXudong Sun, Wenjie Ma, Jiawei Tyler Gu, Zicheng Ma 等OSDI 2024 · 被引用 50 次
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan 等OSDI 2022 · 被引用 44 次
它引用的顶会 Paper4
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- Virtual Consensus in DelosMahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi 等OSDI 2020 · 被引用 42 次
- Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at ScaleShaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis 等OSDI 2020 · 被引用 42 次
相关 Paper
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar 等SOSP 2021 · 被引用 19 次
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- Scaling Large Production Clusters with Partitioned SynchronizationYihui Feng, Zhi Liu, Yunjian Zhao, Tatiana Jin 等USENIX ATC 2021 · 被引用 23 次
- Embracing Imbalance: Dynamic Load Shifting among Microservice Containers in Shared ClustersShutian Luo, Jianxiong Liao, Chenyu Lin, Huanle Xu 等ASPLOS 2025 · 被引用 2 次
- Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User WebsiteUsama Naseer, Luca Niccolini, Udip Pant, Alan Frindell 等SIGCOMM 2020 · 被引用 19 次
