Outage-Watch: Early Prediction of Outages using Extreme Event Regularizer
Shubham Agarwal, Sarthak Chakraborty, Shaddy Garg, Sumit Bisht, Chahat Jain, Ashritha Gonuguntla, Shiv Kumar Saini
Abstract
Cloud services are omnipresent and critical cloud service failure is a fact of life. In order to retain customers and prevent revenue loss, it is important to provide high reliability guarantees for these services. One way to do this is by predicting outages in advance, which can help in reducing the severity as well as time to recovery. It is difficult to forecast critical failures due to the rarity of these events. Moreover, critical failures are ill-defined in terms of observable data. Our proposed method, Outage-Watch, defines critical service outages as deteriorations in the Quality of Service (QoS) captured by a set of metrics. Outage-Watch detects such outages in advance by using current system state to predict whether the QoS metrics will cross a threshold and initiate an extreme event. A mixture of Gaussian is used to model the distribution of the QoS metrics for flexibility and an extreme event regularizer helps in improving learning in tail of the distribution. An outage is predicted if the probability of any one of the QoS metrics crossing threshold changes significantly. Our evaluation on a real-world SaaS company dataset shows that Outage-Watch significantly outperforms traditional methods with an average AUC of 0.98. Additionally, Outage-Watch detects all the outages exhibiting a change in service metrics and reduces the Mean Time To Detection (MTTD) of outages by up to 88% when deployed in an enterprise cloud-service system, demonstrating efficacy of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3af27c01-2f21-4a62-9b38-a24ecabd20bbCited by top-tier papers1
Ask how each one uses itBuilds on11
- TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series DataShreshth Tuli, Giuliano Casale, Nicholas R. JenningsVLDB 2022 · 930 citations
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 262 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao et al.FAST 2020 · 120 citations
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng et al.FSE 2020 · 48 citations
Related papers
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling et al.ASE 2021 · 23 citations
- Maat: Performance Metric Anomaly Anticipation for Cloud Services with Conditional DiffusionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su et al.ASE 2023 · 7 citations
- Outlier-Resilient Web Service QoS PredictionFanghua Ye, Zhiwei Lin, Chuan Chen, Zibin Zheng et al.WWW 2021 · 75 citations
- Monitoring Cloud Service Unreachability at ScaleKapil Agrawal, Viral Mehta, Sundararajan Renganathan, Sreangsu Acharyya et al.INFOCOM 2021 · 1 citation
- ESRO: Experience Assisted Service Reliability against OutagesSarthak Chakraborty, Shubham Agarwal, Shaddy Garg, Abhimanyu Sethia et al.ASE 2023 · 3 citations
