Vapro: performance variance detection and diagnosis for production-run parallel applications
Liyan Zheng, Jidong Zhai, Xiongchao Tang, Haojie Wang, Teng Yu, Yuyang Jin, Shuaiwen Leon Song, Wenguang Chen
Abstract
Performance variance is a serious problem for parallel applications, which can cause performance degradation and make applications' behavior hard to understand. Therefore, detecting and diagnosing performance variance are of crucial importance for users and application developers. However, previous detection approaches either bring too large overhead and hurt applications' performance, or rely on nontrivial source code analysis that is impractical for production-run parallel applications.
In this work, we propose Vapro, a performance variance detection and diagnosis framework for production-run parallel applications. Our approach is based on an important observation that most parallel applications contain code snippets that are repeatedly executed with fixed workload, which can be used for performance variance detection. To effectively identify these snippets at runtime even without program source code, we introduce State Transition Graph (STG) to track program execution and then conduct lightweight workload analysis on STG to locate variance. To diagnose the detected variance, Vapro leverages a progressive diagnosis method based on a hybrid model leveraging variance breakdown and statistical analysis. Results show that the performance overhead of Vapro is only 1.38% on average. Vapro can detect the variance in real applications caused by hardware bugs, memory, and IO. After fixing the detected variance, the standard deviation of the execution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9655cc04-6910-4e5d-a149-f6355d0d9fa9Cited by top-tier papers1
Ask how each one uses itRelated papers
- GVARP: Detecting Performance Variance on Large-Scale Heterogeneous SystemsXin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan et al.SC 2024 · 8 citations
- Effective Performance Issue Diagnosis with Value-Assisted Cost ProfilingLingmei Weng, Yigong Hu, Peng Huang, Jason Nieh et al.EuroSys 2023 · 8 citations
- Using sample-based time series data for automated diagnosis of scalability losses in parallel programsLai Wei, John M. Mellor-CrummeyPPoPP 2020 · 7 citations
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang et al.PPoPP 2022 · 10 citations
- Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsBurak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.SC 2023 · 19 citations
