Outage-Watch: Early Prediction of Outages using Extreme Event Regularizer
Shubham Agarwal, Sarthak Chakraborty, Shaddy Garg, Sumit Bisht, Chahat Jain, Ashritha Gonuguntla, Shiv Kumar Saini
摘要
Cloud services are omnipresent and critical cloud service failure is a fact of life. In order to retain customers and prevent revenue loss, it is important to provide high reliability guarantees for these services. One way to do this is by predicting outages in advance, which can help in reducing the severity as well as time to recovery. It is difficult to forecast critical failures due to the rarity of these events. Moreover, critical failures are ill-defined in terms of observable data. Our proposed method, Outage-Watch, defines critical service outages as deteriorations in the Quality of Service (QoS) captured by a set of metrics. Outage-Watch detects such outages in advance by using current system state to predict whether the QoS metrics will cross a threshold and initiate an extreme event. A mixture of Gaussian is used to model the distribution of the QoS metrics for flexibility and an extreme event regularizer helps in improving learning in tail of the distribution. An outage is predicted if the probability of any one of the QoS metrics crossing threshold changes significantly. Our evaluation on a real-world SaaS company dataset shows that Outage-Watch significantly outperforms traditional methods with an average AUC of 0.98. Additionally, Outage-Watch detects all the outages exhibiting a change in service metrics and reduces the Mean Time To Detection (MTTD) of outages by up to 88% when deployed in an enterprise cloud-service system, demonstrating efficacy of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- TranAD: Deep Transformer Networks for Anomaly Detection in Multivariate Time Series DataShreshth Tuli, Giuliano Casale, Nicholas R. JenningsVLDB 2022 · 被引用 930 次
- DeltaGrad: Rapid retraining of machine learning modelsYinjun Wu, Edgar Dobriban, Susan B. DavidsonICML 2020 · 被引用 262 次
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini 等NeurIPS 2022 · 被引用 185 次
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao 等FAST 2020 · 被引用 120 次
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
相关 Paper
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling 等ASE 2021 · 被引用 23 次
- Maat: Performance Metric Anomaly Anticipation for Cloud Services with Conditional DiffusionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su 等ASE 2023 · 被引用 7 次
- Outlier-Resilient Web Service QoS PredictionFanghua Ye, Zhiwei Lin, Chuan Chen, Zibin Zheng 等WWW 2021 · 被引用 75 次
- Monitoring Cloud Service Unreachability at ScaleKapil Agrawal, Viral Mehta, Sundararajan Renganathan, Sreangsu Acharyya 等INFOCOM 2021 · 被引用 1 次
- ESRO: Experience Assisted Service Reliability against OutagesSarthak Chakraborty, Shubham Agarwal, Shaddy Garg, Abhimanyu Sethia 等ASE 2023 · 被引用 3 次
