On Modular Learning of Distributed Systems for Predicting End-to-End Latency
Chieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang, Zhao Lucis Li, Li Lyna Zhang, Mao Yang, Lidong Zhou
摘要
An emerging trend in cloud deployments is to adopt machine learning (ML) models to characterize end-to-end system performance. Despite early success, such methods can incur significant costs when adapting to the deployment dynamics of distributed systems like service scaling-out and replacement. They require hours or even days for data collection and model training, otherwise models may drift to result in unacceptable inaccuracy. This problem arises from the practice of modeling the entire system with monolithic models. We propose Fluxion, a framework to model end-to-end system latency with modularized learning. Fluxion introduces learning assignment, a new abstraction that allows modeling individual sub-components. With a consistent interface, multiple learning assignments can then be dynamically composed into an inference graph, to model a complex distributed system on the fly. Changes in a system sub-component only involve updating the corresponding learning assignment, thus significantly reducing costs. Using three systems with up to 142 microservices on a 100-VM cluster, Fluxion shows a performance modeling MAE (mean absolute error) up to 68.41% lower than monolithic models. In turn, this lower MAE allows better system performance tuning, e.g., a speed up for 90-percentile end-to-end latency by up to 1.57×. All these are achieved under various system deployment dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On the Feasibility and Benefits of Extensive EvaluationYujie Hui, Miao Yu, Hao Qi, Yifan Gan 等SIGMOD 2025 · 被引用 1 次
- EMA: Efficient Model Adaptation for Learning-based SystemsDaiyang Yu, Xinyu Chen, Yihan Zhang, Yan Liang 等SIGCOMM 2026
它引用的顶会 Paper3
- UDO: Universal Database Optimization using Reinforcement LearningJunxiong Wang, Immanuel Trummer, Debabrota BasuVLDB 2021 · 被引用 53 次
- On the Use of ML for Blackbox System Performance PredictionSilvery Fu, Saurabh Gupta, Radhika Mittal, Sylvia RatnasamyNSDI 2021 · 被引用 43 次
- AutoSys: The Design and Operation of Learning-Augmented SystemsChieh-Jan Mike Liang, Hui Xue, Mao Yang, Lidong Zhou 等USENIX ATC 2020 · 被引用 17 次
相关 Paper
- FastPERT: Towards Fast Microservice Application Latency Prediction via Structural Inductive Bias over PERT NetworksDa Sun Handason Tam, Huanle Xu, Yang Liu, Siyue Xie 等AAAI 2025 · 被引用 5 次
- Sluice: End-to-End Latency Guarantee for Long-running Dataflow SystemsZhaochen She, Yancan Mao, Richard T. B. MaINFOCOM 2026
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- Derm: SLA-aware Resource Management for Highly Dynamic MicroservicesLiao Chen, Shutian Luo, Chenyu Lin, Zizhao Mo 等ISCA 2024 · 被引用 9 次
- Ursa: Lightweight Resource Management for Cloud-Native MicroservicesYanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, Christina DelimitrouHPCA 2024 · 被引用 14 次
