Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud Systems
Jinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang, Zhuangbin Chen, Cong Feng, Zengyin Yang, Yongqiang Yang, Michael R. Lyu
摘要
Ensuring the reliability of cloud systems is critical for both cloud vendors and customers. Cloud systems often rely on virtualization techniques to create instances of hardware resources, such as virtual machines. However, virtualization hinders the observability of cloud systems, making it challenging to diagnose platform-level issues. To improve system observability, we propose to infer functional clusters of instances, i.e., groups of instances having similar functionalities. We first conduct a pilot study on a large-scale cloud system, i.e., Huawei Cloud, demonstrating that instances having similar functionalities share similar communication and resource usage patterns. Motivated by these findings, we formulate the identification of functional clusters as a clustering problem and propose a non-intrusive solution called Prism. Prism adopts a coarse-to-fine clustering strategy. It first partitions instances into coarse-grained chunks based on communication patterns. Within each chunk, Prism further groups instances with similar resource usage patterns to produce fine-grained functional clusters. Such a design reduces noises in the data and allows Prism to process massive instances efficiently. We evaluate Prism on two datasets collected from the real-world production environment of Huawei Cloud. Our experiments show that Prism achieves a v-measure of ∼0.95, surpassing existing state-of-the-art solutions. Additionally, we illustrate the integration of Prism within monitoring systems for enhanced cloud reliability through two real-world use cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen 等ICSE 2025 · 被引用 4 次
- Tracezip: Efficient Distributed Tracing via Trace CompressionZhuangbin Chen, Junsong Pu, Zibin ZhengISSTA 2025 · 被引用 2 次
它引用的顶会 Paper12
- Adaptive Performance Anomaly Detection for Online Service Systems via Pattern SketchingZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang 等ICSE 2022 · 被引用 44 次
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He 等FSE 2020 · 被引用 43 次
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang 等USENIX ATC 2021 · 被引用 43 次
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin 等ASE 2020 · 被引用 33 次
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang 等ASE 2021 · 被引用 29 次
相关 Paper
- CloudCluster: Unearthing the Functional Structure of a Cloud ServiceWeiwu Pang, Sourav Panda, Muhammad Jehangir Amjad, Christophe Diot 等NSDI 2022 · 被引用 3 次
- Analyzing the Communication Clusters in Datacenters✱Klaus-Tycho Foerster, Thibault Marette, Stefan Neumann, Claudia Plant 等WWW 2023 · 被引用 4 次
- Scheduling of Time-Varying Workloads Using Reinforcement LearningShanka Subhra Mondal, Nikhil Sheoran, Subrata MitraAAAI 2021 · 被引用 45 次
- CoFI: Consistency-Guided Fault Injection for Cloud SystemsHaicheng Chen, Wensheng Dou, Dong Wang, Feng QinASE 2020 · 被引用 25 次
- GDPRuler: A Trusted GDPR Monitor for Cloud Data SystemsDimitrios Stavrakakis, Masanori Misono, Julian Pritzi, Harshavardhan Unnibhavi 等CCS 2026
