Predicting Failures of Autoscaling Distributed Applications
Giovanni Denaro, Noura El Moussa, Rahim Heydarov, Francesco Lomio, Mauro Pezzè, Ketai Qiu
Abstract
Predicting failures in production environments allows service providers to activate countermeasures that prevent harming the users of the applications. The most successful approaches predict failures from error states that the current approaches identify from anomalies in time series of fixed sets of KPI values collected at runtime. They cannot handle time series of KPI sets with size that varies over time. Thus these approaches work with applications that run on statically configured sets of components and computational nodes, and do not scale up to the many popular cloud applications that exploit autoscaling. This paper proposes P reface , a novel approach to predict failures in cloud applications that exploit autoscaling. P reface originally augments the neural-network-based failure predictors successfully exploited to predict failures in statically configured applications, with a R ectifier layer that handles KPI sets of highly variable size as the ones collected in cloud autoscaling applications, and reduces those KPIs to a set of rectified-KPIs of fixed size that can be fed to the neural-network predictor. The P reface R ectifier computes the rectified-KPIs as descriptive statistics of the original KPIs, for each logical component of the target application. The descriptive statistics shrink the highly variable sets of KPIs collected at different timestamps to a fixed set of values compatible with the input nodes of the neural-network failure predictor. The neural network can then reveal anomalies that correspond to error states, before they propagate to failures that harm the users of the applications. The experiments on both a commercial application and a widely used academic exemplar confirm that P reface can indeed predict many harmful failures early enough to activate proper countermeasures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95fee7b3-771a-4bd9-a8d5-c6ab55ee3b67Builds on1
Related papers
- PASS: Predictive Auto-Scaling System for Large-scale Enterprise Web ApplicationsYunda Guo, Jiake Ge, Panfeng Guo, Yunpeng Chai et al.WWW 2024 · 10 citations
- Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible InstancesJiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi et al.NSDI 2024 · 54 citations
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling et al.ASE 2021 · 23 citations
- Maat: Performance Metric Anomaly Anticipation for Cloud Services with Conditional DiffusionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su et al.ASE 2023 · 7 citations
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ASE 2021 · 29 citations
