AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck Diagnosis
Bin Dong, Jean Luca Bez, Suren Byna
摘要
Manually diagnosing the I/O performance bottleneck for a single application (hereinafter referred to as the "job level'') is a tedious and error-prone procedure requiring domain scientists to have deep knowledge of complex storage systems. However, existing automatic methods for I/O performance bottleneck diagnosis have one major issue: the granularity of the analysis is at the platform or group level and the diagnosis results cannot be applied to the individual application. To address this issue, we designed and developed a method named "Artificial Intelligence for I/O" (AIIO), which uses AI and its interpretation technology to diagnose I/O performance bottlenecks at the job level automatically. By considering the sparsity of I/O log files, employing multiple AI models for performance prediction, merging diagnosis results across multiple models, and generalizing its performance prediction and diagnosis functions, AIIO can accurately and robustly identify the bottleneck of an even unseen application. Experimental results show that real and unseen applications can use the diagnosis results from AIIO to improve their I/O performance by at most 146 times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- HPC I/O throughput bottleneck analysis with explainable local modelsMihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash 等SC 2020 · 被引用 36 次
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie 等HPDC 2022 · 被引用 26 次
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt 等SC 2021 · 被引用 25 次
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy 等SC 2022 · 被引用 6 次
相关 Paper
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
- Memory-mapped I/O on steroidsAnastasios Papagiannis, Manolis Marazakis, Angelos BilasEuroSys 2021 · 被引用 19 次
- LabStor: A Modular and Extensible Platform for Developing High-Performance, Customized I/O Stacks in UserspaceLuke Logan, Jaime Cernuda Garcia, Jay F. Lofstead, Xian-He Sun 等SC 2022 · 被引用 8 次
- Diagnosing Performance Issues in Application-Defined ResourcesYigong Hu, You-Liang Huang, Haodong Zheng, Yicheng Liu 等OSDI 2026 · 被引用 1 次
