Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the Field
Ronglong Wu, Shuyue Zhou, Jiahao Lu, Zhirong Shen, Zikang Xu, Jiwu Shu, Kunlin Yang, Feilong Lin, Yiming Zhang
Abstract
High-bandwidth memory (HBM) is regarded as a promising technology for fundamentally overcoming the memory wall. It stacks up multiple DRAM dies vertically to dramatically improve the memory access bandwidth. However, this architecture also comes with more severe reliability issues, since HBM not only inherits error patterns of the conventional DRAM, but also introduces new error causes.
In this paper, we conduct the first systematical study on HBM errors, which cover over 460 million error events collected from nineteen data centers and span over two years of deployment under a variety of services. Through error analyses and methodology validations, we confirm that the HBM exhibits different error patterns from conventional DRAM, in terms of spatial locality, temporal correlation, and sensor metrics which make empirical prediction models for DRAM error prediction ineffective for HBM. We design and implement Calchas, a hierarchical failure prediction framework for HBM based on our findings, which integrate spatial, temporal, and sensor information from various device levels to predict upcoming failures. The results demonstrate the feasibility of failure prediction across hierarchical levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85feaa66-c823-4a51-a592-3c47f8b19856Cited by top-tier papers3
- Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsShengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan et al.SC 2025 · 6 citations
- Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory ProtectionJunhwan Kim, Seunghyun Kim, Yesin Ryu, Saeid Gorgin et al.ISCA 2026 · 1 citation
- RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNsHanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae KimISCA 2026
Builds on6
- Exploiting Combined Locality for Wide-Stripe Erasure Coding in Distributed StorageYuchong Hu, Liangfeng Cheng, Qiaori Yao, Patrick P. C. Lee et al.FAST 2021 · 88 citations
- Boosting Full-Node Repair in Erasure-Coded StorageShiyao Lin, Guowen Gong, Zhirong Shen, Patrick P. C. Lee et al.USENIX ATC 2021 · 33 citations
- Cost-aware prediction of uncorrected DRAM errors in the fieldIsaac Boixaderas, Darko Zivanovic, Sergi Moré, Javier Bartolome et al.SC 2020 · 30 citations
- From Correctable Memory Errors to Uncorrectable Memory Errors: What Error Bits TellCong Li, Yu Zhang, Jialei Wang, Hang Chen et al.SC 2022 · 20 citations
- BCD deduplication: effective memory compression using partial cache-line deduplicationSungbo Park, Ingab Kang, Yaebin Moon, Jung Ho Ahn et al.ASPLOS 2021 · 19 citations
Related papers
- Predicting DRAM Failures at Scale: A Two-Stage Approach for Heterogeneous SystemsChenglin Wang, Shouxin Wang, Zhirong Shen, Lu Tang et al.HPCA 2026
- Predicting Future-System Reliability with a Component-Level DRAM Fault ModelJeageun Jung, Mattan ErezMICRO 2023 · 9 citations
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu et al.HPCA 2021 · 9 citations
- HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECCRuizhi Zhu, Yanan Guo, Huize Li, Weidong Cao et al.ISCA 2026
- Characterizing and Mitigating Soft Errors in GPU DRAMMichael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee et al.MICRO 2021 · 45 citations
