SC2024Top-tier venue
MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs
Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu, Yuetsu Kodama
Abstract
Modern High-Performance Computing (HPC) systems play a fundamental role in driving scientific research, as they execute computationally intensive jobs originating from diverse domains. However, HPC jobs are characterized by conflicting computational requirements, which may cause inefficiencies in resource usage, system throughput and energy consumption. One approach to tackling this problem is to distinguish between memory-bound and compute-bound jobs at their submission time, with the goal of making informed decisions about their execution. In this paper, we present MCBound, the first online data-driven framework to classify HPC jobs as memory/compute-bound before job execution, without user intervention. We propose a systematic characterization technique to generate a reference dataset from historical data for initial classification model training. Using the proposed characterization technique, we analyze the data of 2.2 million job runs on the Supercomputer Fugaku1, a production HPC system installed at the RIKEN Center for Computational Science, in Japan. We implement MCBound for Fugaku and classify the jobs executed during February 2024. Our approach is proven effective, as it obtains an F1-macro average score of at least 0.89 as prediction quality, while incurring a negligible overhead on the system’s operations. Our Python-based implementation of MCBound can be seamlessly configured and deployed in other HPC systems.1https://www.fujitsu.com/global/about/innovation/fugaku/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf0de1df-f56e-4a97-8bc6-d833eac47db2Builds on3
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama et al.SC 2020 · 112 citations
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 23 citations
- Prodigy: Towards Unsupervised Anomaly Detection in Production HPC SystemsBurak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz et al.SC 2023 · 19 citations
Related papers
- DPS: Adaptive Power Management for Overprovisioned SystemsJianru Ding, Henry HoffmannSC 2023 · 8 citations
- Job characteristics on large-scale systems: long-term analysis, quantification, and implicationsTirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich et al.SC 2020 · 48 citations
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna et al.HPDC 2020 · 34 citations
- Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsGirish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen et al.OOPSLA 2023 · 4 citations
- A Taxonomy of Error Sources in HPC I/O Machine Learning ModelsMihailo Isakov, Mikaela Currier, Eliakin Del Rosario, Sandeep Madireddy et al.SC 2022 · 6 citations
