On Modular Learning of Distributed Systems for Predicting End-to-End Latency
Chieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang, Zhao Lucis Li, Li Lyna Zhang, Mao Yang, Lidong Zhou
Abstract
An emerging trend in cloud deployments is to adopt machine learning (ML) models to characterize end-to-end system performance. Despite early success, such methods can incur significant costs when adapting to the deployment dynamics of distributed systems like service scaling-out and replacement. They require hours or even days for data collection and model training, otherwise models may drift to result in unacceptable inaccuracy. This problem arises from the practice of modeling the entire system with monolithic models. We propose Fluxion, a framework to model end-to-end system latency with modularized learning. Fluxion introduces learning assignment, a new abstraction that allows modeling individual sub-components. With a consistent interface, multiple learning assignments can then be dynamically composed into an inference graph, to model a complex distributed system on the fly. Changes in a system sub-component only involve updating the corresponding learning assignment, thus significantly reducing costs. Using three systems with up to 142 microservices on a 100-VM cluster, Fluxion shows a performance modeling MAE (mean absolute error) up to 68.41% lower than monolithic models. In turn, this lower MAE allows better system performance tuning, e.g., a speed up for 90-percentile end-to-end latency by up to 1.57×. All these are achieved under various system deployment dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd639c3b-ea57-4ce9-a0aa-e5187d5756fbCited by top-tier papers2
- On the Feasibility and Benefits of Extensive EvaluationYujie Hui, Miao Yu, Hao Qi, Yifan Gan et al.SIGMOD 2025 · 1 citation
- EMA: Efficient Model Adaptation for Learning-based SystemsDaiyang Yu, Xinyu Chen, Yihan Zhang, Yan Liang et al.SIGCOMM 2026
Builds on3
- UDO: Universal Database Optimization using Reinforcement LearningJunxiong Wang, Immanuel Trummer, Debabrota BasuVLDB 2021 · 53 citations
- On the Use of ML for Blackbox System Performance PredictionSilvery Fu, Saurabh Gupta, Radhika Mittal, Sylvia RatnasamyNSDI 2021 · 43 citations
- AutoSys: The Design and Operation of Learning-Augmented SystemsChieh-Jan Mike Liang, Hui Xue, Mao Yang, Lidong Zhou et al.USENIX ATC 2020 · 17 citations
Related papers
- FastPERT: Towards Fast Microservice Application Latency Prediction via Structural Inductive Bias over PERT NetworksDa Sun Handason Tam, Huanle Xu, Yang Liu, Siyue Xie et al.AAAI 2025 · 5 citations
- Sluice: End-to-End Latency Guarantee for Long-running Dataflow SystemsZhaochen She, Yancan Mao, Richard T. B. MaINFOCOM 2026
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- Derm: SLA-aware Resource Management for Highly Dynamic MicroservicesLiao Chen, Shutian Luo, Chenyu Lin, Zizhao Mo et al.ISCA 2024 · 9 citations
- Ursa: Lightweight Resource Management for Cloud-Native MicroservicesYanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, Christina DelimitrouHPCA 2024 · 14 citations
