Monitoring Cloud Service Unreachability at Scale
Kapil Agrawal, Viral Mehta, Sundararajan Renganathan, Sreangsu Acharyya, Venkata N. Padmanabhan, Chakri Kotipalli, Liting Zhao
Abstract
We consider the problem of network unreachability in a global-scale cloud-hosted service that caters to hundreds of millions of users. Even when the service itself is up, the "last mile" between where users are, and the cloud is often the weak link that could render the service unreachable. We present NetDetector, a tool for detecting network-unreachability based on measurements from a client-based HTTP-ping service. NetDetector employs two models. The first, GA (Gaussian Alerts) models temporally averaged raw success rate of the HTTP-pings as a Gaussian distribution and flags significant dips below the mean as unreachability episodes. The second, more sophisticated approach (BB, or Beta-Binomial) models the health of network connectivity as the probability of an access request succeeding, estimates health from noisy samples, and alerts based on dips in health below a client-network-specific SLO (service-level objective) derived from data. These algorithms are enhanced by a drill-down technique that identifies a more precise scope of the unreachability event. We present promising results from GA, which has been in deployment, and the experimental BB detector over a 4-month period. For instance, GA flags 49 country-level unreachability incidents, of which 42 were labelled true positives based on investigation by on-call engineers (OCEs).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8aee32a9-ddd8-4dd5-bafe-3bf35a06ac1bRelated papers
- Outage-Watch: Early Prediction of Outages using Extreme Event RegularizerShubham Agarwal, Sarthak Chakraborty, Shaddy Garg, Sumit Bisht et al.FSE 2023 · 5 citations
- Network Error Logging: Client-side measurement of end-to-end web service reliabilitySam Burnett, Lily Chen, Douglas A. Creager, Misha Efimov et al.NSDI 2020 · 18 citations
- Lightweight Detection of Abnormal Battery Drain Induced by Network Operations of Mobile AppsRun Wang, Marco Brocanelli, Xiaorui WangINFOCOM 2026
- Meaningful AvailabilityTamas Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean et al.NSDI 2020 · 20 citations
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling et al.ASE 2021 · 23 citations
