A Taxonomy of Error Sources in HPC I/O Machine Learning Models
Mihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash, Philip H. Carns, Robert B. Ross, Glenn K. Lockwood, Michel A. Kinsy
摘要
I/O efficiency is crucial to productivity in scientific computing, but the increasing complexity of the system and the applications makes it difficult for practitioners to understand and optimize I/O behavior at scale. Data-driven machine learningbased I/O throughput models offer a solution: they can be used to identify bottlenecks, automate I/O tuning, or optimize job scheduling with minimal human intervention. Unfortunately, current state-of-the-art I/O models are not robust enough for production use and underperform after being deployed.
We analyze multiple years of application, scheduler, and storage system logs on two leadership-class HPC platforms to understand why I/O models underperform in practice. We propose a taxonomy consisting of five categories of I/O modeling errors: poor application and system modeling, inadequate dataset coverage, I/O contention, and I/O noise. We develop litmus tests to quantify each category, allowing researchers to narrow down failure modes, enhance I/O throughput models, and improve future generations of HPC logging and analysis tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
- Neural Ensemble Search for Uncertainty Estimation and Dataset ShiftSheheryar Zaidi, Arber Zela, Thomas Elsken, Chris C. Holmes 等NeurIPS 2021 · 被引用 97 次
- HPC I/O throughput bottleneck analysis with explainable local modelsMihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash 等SC 2020 · 被引用 36 次
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt 等SC 2021 · 被引用 25 次
相关 Paper
- Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsArnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi WangHPDC 2022 · 被引用 12 次
- GLANCED-IO: Taming I/O Optimization for Deep Learning at ScaleRay A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng 等HPDC 2026
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie 等HPDC 2022 · 被引用 26 次
- PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC SystemsRunzhou Han, Suren Byna, Houjun Tang, Bin Dong 等HPDC 2022 · 被引用 9 次
- MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC JobsFrancesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu 等SC 2024 · 被引用 10 次
