FastPERT: Towards Fast Microservice Application Latency Prediction via Structural Inductive Bias over PERT Networks
Da Sun Handason Tam, Huanle Xu, Yang Liu, Siyue Xie, Wing Cheong Lau
Abstract
The recent surge in popularity of cloud-native applications using microservice architectures has led to a focus on accurate end-to-end latency prediction for proactive resource allocation. Existing models leverage Graph Transformers to Microservice Call Graphs or the Program Evaluation and Review Technique (PERT) graphs to capture complex temporal dependencies between microservices. However, these models incur a high computational cost during both training and inference phases. This paper introduces FastPERT, an efficient model for predicting end-to-end latency in microservice applications. FastPERT dissects an execution trace into several microservices tasks, using observations from prior execution traces of the application, akin to the PERT approach. Subsequently, a prediction model is constructed to estimate the completion time for each individual task. This information, coupled with the computational and structural inductive bias of the PERT graph, facilitates the efficient computation of the end-to-end latency of an execution trace. As a result, FastPERT can efficiently capture the complex temporal causality of different microservice tasks without relying on Graph Neural Networks, leading to more accurate and robust latency predictions across a variety of applications.
An evaluation based on datasets generated from large-scale Alibaba microservice traces reveals that FastPERT significantly improves training and inference efficiency without compromising performance, demonstrating its potential as a superior solution for real-time end-to-end latency prediction in cloud-native microservice applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ceb91b0b-d21c-411c-9b4f-56e157f5c59cBuilds on9
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo et al.ASPLOS 2021 · 170 citations
- Erms: Efficient Resource Management for Shared Microservices with SLA GuaranteesShutian Luo, Huanle Xu, Kejiang Ye, Guoyao Xu et al.ASPLOS 2023 · 55 citations
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen et al.USENIX ATC 2023 · 32 citations
Related papers
- PERT-GNN: Latency Prediction for Microservice-based Cloud-Native Applications via Graph Neural NetworksDa Sun Handason Tam, Yang Liu, Huanle Xu, Siyue Xie et al.KDD 2023 · 20 citations
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang et al.NSDI 2023 · 19 citations
- Ursa: Lightweight Resource Management for Cloud-Native MicroservicesYanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, Christina DelimitrouHPCA 2024 · 14 citations
- Slowpoke: End-to-end Throughput Optimization Modeling for Microservice ApplicationsYizheng Xie, Di Jin, Oguzhan Çölkesen, Vasiliki Kalavri et al.NSDI 2026 · 1 citation
- MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsGuangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan et al.WWW 2021 · 152 citations
