SC2020Top-tier venue
Cost-aware prediction of uncorrected DRAM errors in the field
Isaac Boixaderas, Darko Zivanovic, Sergi Moré, Javier Bartolome, David Vicente, Marc Casas, Paul M. Carpenter, Petar Radojkovic, Eduard Ayguadé
Abstract
This paper presents and evaluates a method to predict DRAM uncorrected errors, a leading cause of hardware failures in large-scale HPC clusters. The method uses a random forest classifier, which was trained and evaluated using error logs from two years of production of the MareNostrum 3 supercomputer. By enabling the system to take measures to mitigate node failures, our method reduces lost compute time by up to 57%, a net saving of 21,000 node-hours per year. We release all source code as open source. We also discuss and clarify aspects of methodology that are essential for a DRAM prediction method to be useful in practice. We explain why standard evaluation metrics, such as precision and recall, are insufficient, and base the evaluation on a cost-benefit analysis. This methodology can help ensure that any DRAM error predictor is clear from training bias and has a clear cost-benefit calculation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a52b33b2-f3ea-4de2-a3fc-e5208956e6c5Cited by top-tier papers2
- Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the FieldRonglong Wu, Shuyue Zhou, Jiahao Lu, Zhirong Shen et al.USENIX ATC 2024 · 17 citations
- Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the FieldIsaac Boixaderas, Sergi Moré, Javier Bartolome, David Vicente et al.HPDC 2024 · 1 citation
Related papers
- MemSeer: Leverage Memory Failure Distinctions and Multi-Grained Prediction in Ultra-Scale Heterogeneous X86/ARM ClustersYunfei Gu, Yixuan Liu, Xinyuan Wu, Bo Shao et al.DAC 2025 · 2 citations
- Predicting DRAM Failures at Scale: A Two-Stage Approach for Heterogeneous SystemsChenglin Wang, Shouxin Wang, Zhirong Shen, Lu Tang et al.HPCA 2026
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu et al.HPCA 2021 · 9 citations
- Understanding Memory Failures on a Petascale Arm SystemKurt B. Ferreira, Scott Levy, Joshua Hemmert, Kevin T. PedrettiHPDC 2022 · 8 citations
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao et al.FAST 2020 · 120 citations
