SC2020Top-tier venue
HPC I/O throughput bottleneck analysis with explainable local models
Mihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash, Philip H. Carns, Robert B. Ross, Michel A. Kinsy
Abstract
With the growing complexity of high-performance computing (HPC) systems, achieving high performance can be difficult because of I/O bottlenecks. We analyze multiple years' worth of Darshan logs from the Argonne Leadership Computing Facility's Theta supercomputer in order to understand causes of poor I/O throughput. We present Gauge: a data-driven diagnostic tool for exploring the latent space of supercomputing job features, understanding behaviors of clusters of jobs, and interpreting I/O bottlenecks. We find groups of jobs that at first sight are highly heterogeneous but share certain behaviors, and analyze these groups instead of individual jobs, allowing us to reduce the workload of domain experts and automate I/O performance analysis. We conduct a case study where a system owner using Gauge was able to arrive at several clusters that do not conform to conventional I/O behaviors, as well as find several potential improvements, both on the application level and the system level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb388085-0213-4f9c-b043-debcb9420baeCited by top-tier papers4
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie et al.HPDC 2022 · 26 citations
- Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsArnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi WangHPDC 2022 · 12 citations
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy et al.SC 2022 · 6 citations
- AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck DiagnosisBin Dong, Jean Luca Bez, Suren BynaHPDC 2023 · 5 citations
Related papers
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna et al.HPDC 2020 · 34 citations
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt et al.SC 2021 · 25 citations
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu et al.NSDI 2025 · 3 citations
- DFTracer: An Analysis-Friendly Data Flow Tracer for AI-Driven WorkflowsHariharan Devarajan, Loïc Pottier, Kaushik Velusamy, Huihuo Zheng et al.SC 2024 · 14 citations
- Taming I/O variation on QoS-less HPC storage: what can applications do?Zhenbo Qiao, Qing Liu, Norbert Podhorszki, Scott Klasky et al.SC 2020 · 10 citations
