SC2020Top-tier venue
Job characteristics on large-scale systems: long-term analysis, quantification, and implications
Tirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich, William E. Allcock, Devesh Tiwari
Abstract
HPC workload analysis and resource consumption characteristics are the key to driving better operation practices, system procurement decisions, and designing effective resource management techniques. Unfortunately, the HPC community does not have easy accessibility to long-term introspective work-load analysis and characterization for production-scale HPC systems. This study bridges this gap by providing detailed long-term quantification, characterization, and analysis of job characteristics on two supercomputers: Intrepid and Mira. This study is one of the largest of its kind - covering trends and characteristics for over three billion compute hours, 750 thousand jobs, and spanning a decade. We confirm several long-held conventional wisdom, and identify many previously undiscovered trends and its implications. We also introduce a learning based technique to predict the resource requirement of future jobs with high accuracy, using features available prior to the job submission and without requiring any application-specific tracing or application-intrusive instrumentation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputerWoong Shin, Vladyslav Oles, Ahmad Maroof Karimi, J. Austin Ellis et al.SC 2021 · 41 citations
- DeltaFS: a scalable no-ground-truth filesystem for massively-parallel computingQing Zheng, Charles D. Cranor, Gregory R. Ganger, Garth A. Gibson et al.SC 2021 · 5 citations
Related papers
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna et al.HPDC 2020 · 34 citations
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt et al.SC 2021 · 25 citations
- Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsArnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi WangHPDC 2022 · 12 citations
- Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsGirish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen et al.OOPSLA 2023 · 4 citations
- MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC JobsFrancesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu et al.SC 2024 · 10 citations
