Cooperative Graceful Degradation in Containerized Clouds
Kapil Agrawal, Sangeetha Abdu Jyothi
Abstract
Cloud resilience is crucial for cloud operators and the myriad of applications that rely on the cloud. Today, we lack a mechanism that enables cloud operators to perform graceful degradation of applications while satisfying the application's availability requirements. In this paper, we put forward a vision for automated cloud resilience management with cooperative graceful degradation between applications and cloud operators. First, we investigate techniques for graceful degradation and identify an opportunity for cooperative graceful degradation in public clouds. Second, leveraging criticality tags on containers, we propose diagonal scaling---turning off non-critical containers during capacity crunch scenarios---to maximize the availability of critical services. Third, we design Phoenix, an automated cloud resilience management system that maximizes critical service availability of applications while also considering operator objectives, thereby improving the overall resilience of the infrastructure during failures. We experimentally show that the Phoenix controller running atop Kubernetes can improve critical service availability by up to 2× during large-scale failures. Phoenix can handle failures in a cluster of 100,000 nodes within 10 seconds. We also develop AdaptLab, an open-source resilience benchmarking framework that can emulate realistic cloud environments with real-world application dependency graphs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ace27ea-266a-4fb5-9515-2bb7e0428862Cited by top-tier papers1
Ask how each one uses itBuilds on21
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Network planning with deep reinforcement learningHang Zhu, Varun Gupta, Satyajeet Singh Ahuja, Yuandong Tian et al.SIGCOMM 2021 · 108 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
Related papers
- Phoenix: Detect and Locate Resilience Issues in Blockchain via Context-Sensitive ChaosFuchen Ma, Yuanliang Chen, Yuanhang Zhou, Jingxuan Sun et al.CCS 2023 · 10 citations
- MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination IndexingTianyi Yang, Cheryl Lee, Jiacheng Shen, Yuxin Su et al.ISSTA 2024 · 5 citations
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang et al.SC 2024 · 1 citation
- Optimistic Recovery for High-Availability Software via Partial Process State PreservationYuzhuo Jing, Yuqi Mai, Angting Cai, Yi Chen et al.SOSP 2025 · 1 citation
- Predicting Failures of Autoscaling Distributed ApplicationsGiovanni Denaro, Noura El Moussa, Rahim Heydarov, Francesco Lomio et al.FSE 2024 · 2 citations
