Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems
Arnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi Wang
摘要
Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.
We address such issues by building an end-to-end machine learning (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and generative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- HPC I/O throughput bottleneck analysis with explainable local modelsMihailo Isakov, Eliakin Del Rosario, Sandeep Madireddy, Prasanna Balaprakash 等SC 2020 · 被引用 36 次
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt 等SC 2021 · 被引用 25 次
相关 Paper
- DFTracer: An Analysis-Friendly Data Flow Tracer for AI-Driven WorkflowsHariharan Devarajan, Loïc Pottier, Kaushik Velusamy, Huihuo Zheng 等SC 2024 · 被引用 14 次
- Bringing Differential Privacy to HPC: Privacy-Preserving Transformations of HPC TracesAna Luisa Veroneze Solórzano, Rohan Basu Roy, Benjamin Schwaller, Sara Petra Walton 等HPDC 2025
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy 等SC 2022 · 被引用 6 次
- Job characteristics on large-scale systems: long-term analysis, quantification, and implicationsTirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich 等SC 2020 · 被引用 48 次
