HPC I/O throughput bottleneck analysis with explainable local models
Mihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash, Philip H. Carns, Robert B. Ross, Michel A. Kinsy
摘要
With the growing complexity of high-performance computing (HPC) systems, achieving high performance can be difficult because of I/O bottlenecks. We analyze multiple years' worth of Darshan logs from the Argonne Leadership Computing Facility's Theta supercomputer in order to understand causes of poor I/O throughput. We present Gauge: a data-driven diagnostic tool for exploring the latent space of supercomputing job features, understanding behaviors of clusters of jobs, and interpreting I/O bottlenecks. We find groups of jobs that at first sight are highly heterogeneous but share certain behaviors, and analyze these groups instead of individual jobs, allowing us to reduce the workload of domain experts and automate I/O performance analysis. We conduct a case study where a system owner using Gauge was able to arrive at several clusters that do not conform to conventional I/O behaviors, as well as find several potential improvements, both on the application level and the system level.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie 等HPDC 2022 · 被引用 26 次
- Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsArnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi WangHPDC 2022 · 被引用 12 次
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy 等SC 2022 · 被引用 6 次
- AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck DiagnosisBin Dong, Jean Luca Bez, Suren BynaHPDC 2023 · 被引用 5 次
相关 Paper
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt 等SC 2021 · 被引用 25 次
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu 等NSDI 2025 · 被引用 3 次
- DFTracer: An Analysis-Friendly Data Flow Tracer for AI-Driven WorkflowsHariharan Devarajan, Loïc Pottier, Kaushik Velusamy, Huihuo Zheng 等SC 2024 · 被引用 14 次
- Taming I/O variation on QoS-less HPC storage: what can applications do?Zhenbo Qiao, Qing Liu, Norbert Podhorszki, Scott Klasky 等SC 2020 · 被引用 10 次
