CAPA: An Architecture For Operating Cluster Networks With High Availability
Bingzhe Liu, Colin Scott, Mukarram Tariq, Andrew D. Ferguson, Phillipa Gill, Richard Alimi, Omid Alipourfard, Deepak Arulkannan, Virginia Beauregard, Patrick Conner, Philip Brighten Godfrey, Xander Lin
摘要
Management operations are a major source of outages for networks. A number of best practices designed to reduce and mitigate such outages are well known, but their enforcement has been challenging, leaving the network vulnerable to inadvertent mistakes and gaps which repeatedly result in outages. We present our experiences with CAPA, Google's "containment and prevention architecture" for regulating management operations on our cluster networking fleet. Our goal with CAPA is to limit the systems where strict adherence to best practices is required, so that availability of the network is not dependent on the good intentions of every engineer and operator. We enumerate the features of CAPA which we have found to be necessary to effectively enforce best practices within a thin "regulation" layer. We evaluate CAPA based on case studies of outages prevented, counter-factual analysis of past incidents, and known limitations. Management-planerelated outages have substantially reduced both in frequency and severity, with a 82% reduction in cumulative duration of incidents normalized to fleet size over five years.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh 等SIGCOMM 2022 · 被引用 230 次
- Plankton: Scalable network configuration verification through model checkingSanthosh Prabhu, Kuan-Yen Chou, Ali Kheradmand, Brighten Godfrey 等NSDI 2020 · 被引用 130 次
- Orion: Google's Software-Defined Networking Control PlaneAndrew D. Ferguson, Steve D. Gribble, Chi-Yao Hong, Charles Killian 等NSDI 2021 · 被引用 95 次
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu 等SIGCOMM 2022 · 被引用 82 次
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan 等SIGCOMM 2021 · 被引用 63 次
相关 Paper
- Improving Network Availability with Protective ReRouteDavid Wetherall, Abdul Kabbani, Van Jacobson, Jim Winget 等SIGCOMM 2023 · 被引用 18 次
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin 等EuroSys 2021 · 被引用 60 次
- Fathom: Understanding Datacenter Application Network PerformanceMubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh 等SIGCOMM 2023 · 被引用 7 次
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu 等NSDI 2025 · 被引用 3 次
- Concord: Learning Network Configuration ContractsRyan Beckett, Francis Y. Yan, Raghunadha Reddy Pocha, Vineesh V. Raj 等EuroSys 2026
