MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs
Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu, Yuetsu Kodama
摘要
Modern High-Performance Computing (HPC) systems play a fundamental role in driving scientific research, as they execute computationally intensive jobs originating from diverse domains. However, HPC jobs are characterized by conflicting computational requirements, which may cause inefficiencies in resource usage, system throughput and energy consumption. One approach to tackling this problem is to distinguish between memory-bound and compute-bound jobs at their submission time, with the goal of making informed decisions about their execution. In this paper, we present MCBound, the first online data-driven framework to classify HPC jobs as memory/compute-bound before job execution, without user intervention. We propose a systematic characterization technique to generate a reference dataset from historical data for initial classification model training. Using the proposed characterization technique, we analyze the data of 2.2 million job runs on the Supercomputer Fugaku1, a production HPC system installed at the RIKEN Center for Computational Science, in Japan. We implement MCBound for Fugaku and classify the jobs executed during February 2024. Our approach is proven effective, as it obtains an F1-macro average score of at least 0.89 as prediction quality, while incurring a negligible overhead on the system’s operations. Our Python-based implementation of MCBound can be seamlessly configured and deployed in other HPC systems.1https://www.fujitsu.com/global/about/innovation/fugaku/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama 等SC 2020 · 被引用 112 次
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 被引用 23 次
- Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsBurak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz 等SC 2023 · 被引用 19 次
相关 Paper
- DPS: Adaptive Power Management for Overprovisioned SystemsJianru Ding, Henry HoffmannSC 2023 · 被引用 8 次
- Job characteristics on large-scale systems: long-term analysis, quantification, and implicationsTirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich 等SC 2020 · 被引用 48 次
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsGirish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen 等OOPSLA 2023 · 被引用 4 次
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy 等SC 2022 · 被引用 6 次
