FastPERT: Towards Fast Microservice Application Latency Prediction via Structural Inductive Bias over PERT Networks
Da Sun Handason Tam, Huanle Xu, Yang Liu, Siyue Xie, Wing Cheong Lau
摘要
The recent surge in popularity of cloud-native applications using microservice architectures has led to a focus on accurate end-to-end latency prediction for proactive resource allocation. Existing models leverage Graph Transformers to Microservice Call Graphs or the Program Evaluation and Review Technique (PERT) graphs to capture complex temporal dependencies between microservices. However, these models incur a high computational cost during both training and inference phases. This paper introduces FastPERT, an efficient model for predicting end-to-end latency in microservice applications. FastPERT dissects an execution trace into several microservices tasks, using observations from prior execution traces of the application, akin to the PERT approach. Subsequently, a prediction model is constructed to estimate the completion time for each individual task. This information, coupled with the computational and structural inductive bias of the PERT graph, facilitates the efficient computation of the end-to-end latency of an execution trace. As a result, FastPERT can efficiently capture the complex temporal causality of different microservice tasks without relying on Graph Neural Networks, leading to more accurate and robust latency predictions across a variety of applications.
An evaluation based on datasets generated from large-scale Alibaba microservice traces reveals that FastPERT significantly improves training and inference efficiency without compromising performance, demonstrating its potential as a superior solution for real-time end-to-end latency prediction in cloud-native microservice applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- Erms: Efficient Resource Management for Shared Microservices with SLA GuaranteesShutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu 等ASPLOS 2023 · 被引用 55 次
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen 等USENIX ATC 2023 · 被引用 32 次
相关 Paper
- PERT-GNN: Latency Prediction for Microservice-based Cloud-Native Applications via Graph Neural NetworksDa Sun Handason Tam, Yang Liu, Huanle Xu, Siyue Xie 等KDD 2023 · 被引用 20 次
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang 等NSDI 2023 · 被引用 19 次
- Ursa: Lightweight Resource Management for Cloud-Native MicroservicesYanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, Christina DelimitrouHPCA 2024 · 被引用 14 次
- Slowpoke: End-to-end Throughput Optimization Modeling for Microservice ApplicationsYizheng Xie, Di Jin, Oguzhan Çölkesen, Vasiliki Kalavri 等NSDI 2026 · 被引用 1 次
- MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsGuangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan 等WWW 2021 · 被引用 152 次
