USENIX ATC2024顶会
Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the Field
Ronglong Wu, Shuyue Zhou, Jiahao Lu, Zhirong Shen, Zikang Xu, Jiwu Shu, Kunlin Yang, Feilong Lin, Yiming Zhang
摘要
High-bandwidth memory (HBM) is regarded as a promising technology for fundamentally overcoming the memory wall. It stacks up multiple DRAM dies vertically to dramatically improve the memory access bandwidth. However, this architecture also comes with more severe reliability issues, since HBM not only inherits error patterns of the conventional DRAM, but also introduces new error causes.
In this paper, we conduct the first systematical study on HBM errors, which cover over 460 million error events collected from nineteen data centers and span over two years of deployment under a variety of services. Through error analyses and methodology validations, we confirm that the HBM exhibits different error patterns from conventional DRAM, in terms of spatial locality, temporal correlation, and sensor metrics which make empirical prediction models for DRAM error prediction ineffective for HBM. We design and implement Calchas, a hierarchical failure prediction framework for HBM based on our findings, which integrate spatial, temporal, and sensor information from various device levels to predict upcoming failures. The results demonstrate the feasibility of failure prediction across hierarchical levels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsShengkun Cui, Archit Patke, Hung Nguyen, Aditya Ranjan 等SC 2025 · 被引用 6 次
- Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory ProtectionJunhwan Kim, Seunghyun Kim, Yesin Ryu, Saeid Gorgin 等ISCA 2026 · 被引用 1 次
- RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNsHanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae KimISCA 2026
它引用的顶会 Paper6
- Exploiting Combined Locality for Wide-Stripe Erasure Coding in Distributed StorageYuchong Hu, Liangfeng Cheng, Qiaori Yao, Patrick P. C. Lee 等FAST 2021 · 被引用 88 次
- Boosting Full-Node Repair in Erasure-Coded StorageShiyao Lin, Guowen Gong, Zhirong Shen, Patrick P. C. Lee 等USENIX ATC 2021 · 被引用 33 次
- Cost-aware prediction of uncorrected DRAM errors in the fieldIsaac Boixaderas, Darko Zivanovic, Sergi Moré, Javier Bartolome 等SC 2020 · 被引用 30 次
- From Correctable Memory Errors to Uncorrectable Memory Errors: What Error Bits TellCong Li, Yu Zhang, Jialei Wang, Hang Chen 等SC 2022 · 被引用 20 次
- BCD deduplication: effective memory compression using partial cache-line deduplicationSungbo Park, Ingab Kang, Yaebin Moon, Jung Ho Ahn 等ASPLOS 2021 · 被引用 19 次
相关 Paper
- Predicting DRAM Failures at Scale: A Two-Stage Approach for Heterogeneous SystemsChenglin Wang, Shouxin Wang, Zhirong Shen, Lu Tang 等HPCA 2026
- Predicting Future-System Reliability with a Component-Level DRAM Fault ModelJeageun Jung, Mattan ErezMICRO 2023 · 被引用 9 次
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu 等HPCA 2021 · 被引用 9 次
- HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECCRuizhi Zhu, Yanan Guo, Huize Li, Weidong Cao 等ISCA 2026
- Characterizing and Mitigating Soft Errors in GPU DRAMMichael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee 等MICRO 2021 · 被引用 45 次
