Cooperative Graceful Degradation in Containerized Clouds
Kapil Agrawal, Sangeetha Abdu Jyothi
摘要
Cloud resilience is crucial for cloud operators and the myriad of applications that rely on the cloud. Today, we lack a mechanism that enables cloud operators to perform graceful degradation of applications while satisfying the application's availability requirements. In this paper, we put forward a vision for automated cloud resilience management with cooperative graceful degradation between applications and cloud operators. First, we investigate techniques for graceful degradation and identify an opportunity for cooperative graceful degradation in public clouds. Second, leveraging criticality tags on containers, we propose diagonal scaling---turning off non-critical containers during capacity crunch scenarios---to maximize the availability of critical services. Third, we design Phoenix, an automated cloud resilience management system that maximizes critical service availability of applications while also considering operator objectives, thereby improving the overall resilience of the infrastructure during failures. We experimentally show that the Phoenix controller running atop Kubernetes can improve critical service availability by up to 2× during large-scale failures. Phoenix can handle failures in a cluster of 100,000 nodes within 10 seconds. We also develop AdaptLab, an open-source resilience benchmarking framework that can emulate realistic cloud environments with real-world application dependency graphs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- Network planning with deep reinforcement learningHang Zhu, Varun Gupta, Satyajeet Singh Ahuja, Yuandong Tian 等SIGCOMM 2021 · 被引用 108 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
相关 Paper
- Phoenix: Detect and Locate Resilience Issues in Blockchain via Context-Sensitive ChaosFuchen Ma, Yuanliang Chen, Yuanhang Zhou, Jingxuan Sun 等CCS 2023 · 被引用 10 次
- MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination IndexingTianyi Yang, Cheryl Lee, Jiacheng Shen, Yuxin Su 等ISSTA 2024 · 被引用 5 次
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang 等SC 2024 · 被引用 1 次
- Optimistic Recovery for High-Availability Software via Partial Process State PreservationYuzhuo Jing, Yuqi Mai, Angting Cai, Yi Chen 等SOSP 2025 · 被引用 1 次
- Predicting Failures of Autoscaling Distributed ApplicationsGiovanni Denaro, Noura El Moussa, Rahim Heydarov, Francesco Lomio 等FSE 2024 · 被引用 2 次
