Job characteristics on large-scale systems: long-term analysis, quantification, and implications
Tirthak Patel, Zhengchun Liu, Raj Kettimuthu, Paul Rich, William E. Allcock, Devesh Tiwari
摘要
HPC workload analysis and resource consumption characteristics are the key to driving better operation practices, system procurement decisions, and designing effective resource management techniques. Unfortunately, the HPC community does not have easy accessibility to long-term introspective work-load analysis and characterization for production-scale HPC systems. This study bridges this gap by providing detailed long-term quantification, characterization, and analysis of job characteristics on two supercomputers: Intrepid and Mira. This study is one of the largest of its kind - covering trends and characteristics for over three billion compute hours, 750 thousand jobs, and spanning a decade. We confirm several long-held conventional wisdom, and identify many previously undiscovered trends and its implications. We also introduce a learning based technique to predict the resource requirement of future jobs with high accuracy, using features available prior to the job submission and without requiring any application-specific tracing or application-intrusive instrumentation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen 等SC 2021 · 被引用 136 次
- Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputerWoong Shin, Vladyslav Oles, Ahmad Maroof Karimi, J. Austin Ellis 等SC 2021 · 被引用 41 次
- DeltaFS: a scalable no-ground-truth filesystem for massively-parallel computingQing Zheng, Charles D. Cranor, Gregory R. Ganger, Garth A. Gibson 等SC 2021 · 被引用 5 次
相关 Paper
- Towards HPC I/O Performance Prediction through Large-scale Log AnalysisSunggon Kim, Alex Sim, Kesheng Wu, Suren Byna 等HPDC 2020 · 被引用 34 次
- Systematically inferring I/O performance variability by examining repetitive job behaviorEmily Costa, Tirthak Patel, Benjamin Schwaller, Jim M. Brandt 等SC 2021 · 被引用 25 次
- Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsArnab K. Paul, Jong Youl Choi, Ahmad Maroof Karimi, Feiyi WangHPDC 2022 · 被引用 12 次
- Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsGirish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen 等OOPSLA 2023 · 被引用 4 次
- MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC JobsFrancesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu 等SC 2024 · 被引用 10 次
