Understanding Memory Failures on a Petascale Arm System
Kurt B. Ferreira, Scott Levy, Joshua Hemmert, Kevin T. Pedretti
Abstract
New and novel HPC platforms provide interesting challenges and opportunities. Analysis of these systems can provide a better understanding of both the specific platform being studied as well as large-scale systems in general. Arm is one such architecture that has been explored in HPC for several years, however little is still known about its viability for supporting large-scale production workloads in terms of system reliability. The Astra system at Sandia National Laboratories was the first public peta-FLOPS Arm-based system on the Top500 and has been successfully running production HPC applications for a couple of years. In this paper, we analyze memory failure data collected from Astra while the system was in production running unclassified applications. This analysis revealed several interesting contributions related to both the Arm platform and to HPC systems in general. First, we outline the number of components replaced due to reliability issues in standing-up this first-of-its-kind, large-scale HPC system. We show the distribution differences between correctable DRAM faults and errors on Astra, showing that, not properly accounting for faults can lead to erroneous conclusions. Additionally, we characterize DRAM faults on the system and show contrary to existing work that memory faults are uniformly distributed across CPU socket, DRAM column, bank and rack region, but are not uniform across node, DIMM rank, DIMM slot on the motherboard, and system rack: some racks, ranks and DIMM slots experience more faults than others. Similarly, we show the impact of temperature and power on DRAM correctable errors. Finally, we make a detailed comparison of results presented here with the positional affects found in several previous large-scale reliability studies. The results of this analysis provide valuable guidance to organizations standing-up first-in- class platforms in HPC, organizations using Arm in HPC, and the entire large-scale HPC community in general.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe3123ea-704f-4f9f-bd9b-ae09abf01ec8Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama et al.SC 2020 · 112 citations
- GPU lifetimes on titan supercomputer: survival analysis and reliabilityGeorge Ostrouchov, Don Maxwell, Rizwan A. Ashraf, Christian Engelmann et al.SC 2020 · 38 citations
- Chronicles of astra: challenges and lessons from the first petascale arm supercomputerKevin T. Pedretti, Andrew J. Younge, Simon D. Hammond, James H. Laros III et al.SC 2020 · 13 citations
Related papers
- DrCCTProf: a fine-grained call path profiler for ARM-based clustersQidong Zhao, Xu Liu, Milind ChabbiSC 2020 · 11 citations
- Predicting DRAM Failures at Scale: A Two-Stage Approach for Heterogeneous SystemsChenglin Wang, Shouxin Wang, Zhirong Shen, Lu Tang et al.HPCA 2026
- Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsShengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan et al.SC 2025 · 6 citations
- Cost-aware prediction of uncorrected DRAM errors in the fieldIsaac Boixaderas, Darko Zivanovic, Sergi Moré, Javier Bartolome et al.SC 2020 · 30 citations
- MemSeer: Leverage Memory Failure Distinctions and Multi-Grained Prediction in Ultra-Scale Heterogeneous X86/ARM ClustersYunfei Gu, Yixuan Liu, Xinyuan Wu, Bo Shao et al.DAC 2025 · 2 citations
