On the Use of ML for Blackbox System Performance Prediction
Silvery Fu, Saurabh Gupta, Radhika Mittal, Sylvia Ratnasamy
Abstract
There is a growing body of work that reports positive results from applying ML-based performance prediction to a particular application or use-case (e.g., server configuration, capacity planning). Yet, a critical question remains unanswered: does ML make prediction simpler (i.e., allowing us to treat systems as blackboxes) and general (i.e., across a range of applications and use-cases)? After all, the potential for simplicity and generality is a key part of what makes ML-based prediction so attractive compared to the traditional approach of relying on handcrafted and specialized performance models. In this paper, we attempt to answer this broader question. We develop a methodology for systematically diagnosing whether, when, and why ML does (not) work for performance prediction, and identify steps to improve predictability.
We apply our methodology to test 6 ML models in predicting the performance of 13 real-world applications. We find that 12 out of our 13 applications exhibit inherent variability in performance that fundamentally limits prediction accuracy. Our findings motivate the need for system-level modifications and/or ML-level extensions that can improve predictability, showing how ML fails to be an easy-to-use predictor. On implementing and evaluating these changes, we find that while they do improve the overall prediction accuracy, prediction error remains high for multiple realistic scenarios, showing how ML fails as a general predictor. Hence our answer is clear: ML is not a general and easy-to-use hammer for system performance prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f02182f9-e1ee-4ff3-9730-f2238672f53aCited by top-tier papers7
- Automated SmartNIC Offloading Insights for Network FunctionsYiming Qiu, Jiarong Xing, Kuo-Feng Hsu, Qiao Kang et al.SOSP 2021 · 37 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang et al.NSDI 2023 · 19 citations
- Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data ApplicationsHani Al-Sayeh, Bunjamin Memishi, Muhammad Attahir Jibril, Marcus Paradies et al.SIGMOD 2022 · 16 citations
- On the Feasibility and Benefits of Extensive EvaluationYujie Hui, Miao Yu, Hao Qi, Yifan Gan et al.SIGMOD 2025 · 1 citation
Builds on1
Related papers
- Applied Online Algorithms with Heterogeneous PredictorsJessica Maghakian, Russell Lee, Mohammad Hajiesmaili, Jian Li et al.ICML 2023 · 7 citations
- Towards Improving the Trustworthiness of Hardware based Malware Detector using Online Uncertainty EstimationHarshit Kumar, Nikhil Chawla, Saibal MukhopadhyayDAC 2021 · 6 citations
- Analysing the Impact of Workloads on Modeling the Performance of Configurable Software SystemsStefan Mühlbauer, Florian Sattler, Christian Kaltenecker, Johannes Dorn et al.ICSE 2023 · 20 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance PredictionKaixuan Zhang, Yunfan Cui, Shuhao Zhang, Chutong Ding et al.ISCA 2026 · 2 citations
