Using sample-based time series data for automated diagnosis of scalability losses in parallel programs
Lai Wei, John M. Mellor-Crummey
Abstract
The performance of many parallel applications has failed to scale as fast as successive generations of hardware on which these applications execute. To understand the cause of scalability losses, experts use performance tools to monitor and analyze application behavior. Profiles generated by performance tools can usually indicate the presence of scalability losses while time series data are generally necessary to pinpoint the root causes of such losses. However, manual analysis of time series data can be difficult in executions with a large number of processes, long running times, and deep call chains. This paper describes an automated framework that analyzes sample-based time series data to diagnose scalability losses in parallel executions. The framework's automated diagnosis of scalability losses indicates their symptoms, severity, and causes. Two case studies illustrate the effectiveness of this framework. When compared to a tool that analyzes performance using instrumentation-based traces, our overhead for collecting sample-based time series is 1/28 in time and 1/1600 in space while our automated analysis takes 1/25 of the time.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cd51dbda-b6a7-473a-90bb-c72b13d8cc64Related papers
- ScalAna: automating scaling loss detection with graph analysisYuyang Jin, Haojie Wang, Teng Yu, Xiongchao Tang et al.SC 2020 · 16 citations
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang et al.PPoPP 2022 · 10 citations
- Extracting clean performance models from tainted programsMarcin Copik, Alexandru Calotoiu, Tobias Grosser, Nicolas Wicki et al.PPoPP 2021 · 16 citations
- Vapro: performance variance detection and diagnosis for production-run parallel applicationsLiyan Zheng, Jidong Zhai, Xiongchao Tang, Haojie Wang et al.PPoPP 2022 · 15 citations
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh et al.NSDI 2024 · 4 citations
