On the Use of ML for Blackbox System Performance Prediction
Silvery Fu, Saurabh Gupta, Radhika Mittal, Sylvia Ratnasamy
摘要
There is a growing body of work that reports positive results from applying ML-based performance prediction to a particular application or use-case (e.g., server configuration, capacity planning). Yet, a critical question remains unanswered: does ML make prediction simpler (i.e., allowing us to treat systems as blackboxes) and general (i.e., across a range of applications and use-cases)? After all, the potential for simplicity and generality is a key part of what makes ML-based prediction so attractive compared to the traditional approach of relying on handcrafted and specialized performance models. In this paper, we attempt to answer this broader question. We develop a methodology for systematically diagnosing whether, when, and why ML does (not) work for performance prediction, and identify steps to improve predictability.
We apply our methodology to test 6 ML models in predicting the performance of 13 real-world applications. We find that 12 out of our 13 applications exhibit inherent variability in performance that fundamentally limits prediction accuracy. Our findings motivate the need for system-level modifications and/or ML-level extensions that can improve predictability, showing how ML fails to be an easy-to-use predictor. On implementing and evaluating these changes, we find that while they do improve the overall prediction accuracy, prediction error remains high for multiple realistic scenarios, showing how ML fails as a general predictor. Hence our answer is clear: ML is not a general and easy-to-use hammer for system performance prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Automated SmartNIC Offloading Insights for Network FunctionsYiming Qiu, Jiarong Xing, Kuo-Feng Hsu, Qiao Kang 等SOSP 2021 · 被引用 37 次
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 等SOSP 2023 · 被引用 31 次
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang 等NSDI 2023 · 被引用 19 次
- Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data ApplicationsHani Al-Sayeh, Bunjamin Memishi, Muhammad Attahir Jibril, Marcus Paradies 等SIGMOD 2022 · 被引用 16 次
- On the Feasibility and Benefits of Extensive EvaluationYujie Hui, Miao Yu, Hao Qi, Yifan Gan 等SIGMOD 2025 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Applied Online Algorithms with Heterogeneous PredictorsJessica Maghakian, Russell Lee, Mohammad Hajiesmaili, Jian Li 等ICML 2023 · 被引用 7 次
- Towards Improving the Trustworthiness of Hardware based Malware Detector using Online Uncertainty EstimationHarshit Kumar, Nikhil Chawla, Saibal MukhopadhyayDAC 2021 · 被引用 6 次
- Analysing the Impact of Workloads on Modeling the Performance of Configurable Software SystemsStefan Mühlbauer, Florian Sattler, Christian Kaltenecker, Johannes Dorn 等ICSE 2023 · 被引用 20 次
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance PredictionKaixuan Zhang, Yunfan Cui, Shuhao Zhang, Chutong Ding 等ISCA 2026 · 被引用 2 次
