Vapro: performance variance detection and diagnosis for production-run parallel applications
Liyan Zheng, Jidong Zhai, Xiongchao Tang, Haojie Wang, Teng Yu, Yuyang Jin, Shuaiwen Leon Song, Wenguang Chen
摘要
Performance variance is a serious problem for parallel applications, which can cause performance degradation and make applications' behavior hard to understand. Therefore, detecting and diagnosing performance variance are of crucial importance for users and application developers. However, previous detection approaches either bring too large overhead and hurt applications' performance, or rely on nontrivial source code analysis that is impractical for production-run parallel applications.
In this work, we propose Vapro, a performance variance detection and diagnosis framework for production-run parallel applications. Our approach is based on an important observation that most parallel applications contain code snippets that are repeatedly executed with fixed workload, which can be used for performance variance detection. To effectively identify these snippets at runtime even without program source code, we introduce State Transition Graph (STG) to track program execution and then conduct lightweight workload analysis on STG to locate variance. To diagnose the detected variance, Vapro leverages a progressive diagnosis method based on a hybrid model leveraging variance breakdown and statistical analysis. Results show that the performance overhead of Vapro is only 1.38% on average. Vapro can detect the variance in real applications caused by hardware bugs, memory, and IO. After fixing the detected variance, the standard deviation of the execution
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- GVARP: Detecting Performance Variance on Large-Scale Heterogeneous SystemsXin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan 等SC 2024 · 被引用 8 次
- Effective Performance Issue Diagnosis with Value-Assisted Cost ProfilingLingmei Weng, Yigong Hu, Peng Huang, Jason Nieh 等EuroSys 2023 · 被引用 8 次
- Using sample-based time series data for automated diagnosis of scalability losses in parallel programsLai Wei, John M. Mellor-CrummeyPPoPP 2020 · 被引用 7 次
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang 等PPoPP 2022 · 被引用 10 次
- Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsBurak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz 等SC 2023 · 被引用 19 次
