Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud Systems
Jinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang, Zhuangbin Chen, Cong Feng, Zengyin Yang, Yongqiang Yang, Michael R. Lyu
Abstract
Ensuring the reliability of cloud systems is critical for both cloud vendors and customers. Cloud systems often rely on virtualization techniques to create instances of hardware resources, such as virtual machines. However, virtualization hinders the observability of cloud systems, making it challenging to diagnose platform-level issues. To improve system observability, we propose to infer functional clusters of instances, i.e., groups of instances having similar functionalities. We first conduct a pilot study on a large-scale cloud system, i.e., Huawei Cloud, demonstrating that instances having similar functionalities share similar communication and resource usage patterns. Motivated by these findings, we formulate the identification of functional clusters as a clustering problem and propose a non-intrusive solution called Prism. Prism adopts a coarse-to-fine clustering strategy. It first partitions instances into coarse-grained chunks based on communication patterns. Within each chunk, Prism further groups instances with similar resource usage patterns to produce fine-grained functional clusters. Such a design reduces noises in the data and allows Prism to process massive instances efficiently. We evaluate Prism on two datasets collected from the real-world production environment of Huawei Cloud. Our experiments show that Prism achieves a v-measure of ∼0.95, surpassing existing state-of-the-art solutions. Additionally, we illustrate the integration of Prism within monitoring systems for enhanced cloud reliability through two real-world use cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen et al.ICSE 2025 · 4 citations
- Tracezip: Efficient Distributed Tracing via Trace CompressionZhuangbin Chen, Junsong Pu, Zibin ZhengISSTA 2025 · 2 citations
Builds on12
- Adaptive Performance Anomaly Detection for Online Service Systems via Pattern SketchingZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ICSE 2022 · 44 citations
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He et al.FSE 2020 · 43 citations
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang et al.USENIX ATC 2021 · 43 citations
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin et al.ASE 2020 · 33 citations
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ASE 2021 · 29 citations
Related papers
- CloudCluster: Unearthing the Functional Structure of a Cloud ServiceWeiwu Pang, Sourav Panda, Muhammad Jehangir Amjad, Christophe Diot et al.NSDI 2022 · 3 citations
- Analyzing the Communication Clusters in Datacenters✱Klaus-Tycho Foerster, Thibault Marette, Stefan Neumann, Claudia Plant et al.WWW 2023 · 4 citations
- Scheduling of Time-Varying Workloads Using Reinforcement LearningShanka Subhra Mondal, Nikhil Sheoran, Subrata MitraAAAI 2021 · 45 citations
- CoFI: Consistency-Guided Fault Injection for Cloud SystemsHaicheng Chen, Wensheng Dou, Dong Wang, Feng QinASE 2020 · 25 citations
- GDPRuler: A Trusted GDPR Monitor for Cloud Data SystemsDimitrios Stavrakakis, Masanori Misono, Julian Pritzi, Harshavardhan Unnibhavi et al.CCS 2026
