DRAM Fault Classification through Large-Scale Field Monitoring for Robust Memory RAS Management
Hoiju Chung, Euisang Oh, Seungmin Baek, Hyeongshin Yoon, Jaesung Yoo, Sanghwan Lee, Yongjun Lee, Arhatha Bramhanand, Brett Dodds, Yang Zhou, Nam Sung Kim
Abstract
As DRAM technology scales down, maintaining prior levels of reliability becomes increasingly challenging due to heightened susceptibility to faults.This growing concern underscores the need for effective in-field fault monitoring and management.Addressing this challenge, this work first introduces a refined DRAM fault classification, derived from the correlation between DRAM addresses and their underlying architectural hierarchies of faulty DRAMs.Building on this classification, we propose a comprehensive memory fault management strategy adequate to each DRAM fault.We further formalize and implement a structured remediation framework, enabling fault-specific mitigation.Utilizing this methodology, we conduct a large-scale field study on DDR4 x4 and DDR5 10x4 RDIMMs, uncovering that over 98% of DDR4 x4 RDIMM faults are tightly coupled with intra-bank architectural characteristics.In particular, we identify that early detection and localized management of faults within 2×2 MATs 1 -encompassing the vast majority of clustered row faults-are crucial to achieving resilient memory subsystems.We also observe intrinsic correlations among faulty Row Addresses (RAs), indicating unavoidable structural address 1 More rigorous definition on MAT will be discussed in §4.1.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6a0eb9a1-7c2d-4967-9eb1-101fc19849d7Cited by top-tier papers2
- Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory ProtectionJunhwan Kim, Seunghyun Kim, Yesin Ryu, Saeid Gorgin et al.ISCA 2026 · 1 citation
- RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNsHanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae KimISCA 2026
Related papers
- Predicting DRAM Failures at Scale: A Two-Stage Approach for Heterogeneous SystemsChenglin Wang, Shouxin Wang, Zhirong Shen, Lu Tang et al.HPCA 2026
- Predicting Future-System Reliability with a Component-Level DRAM Fault ModelJeageun Jung, Mattan ErezMICRO 2023 · 9 citations
- TRRespass: Exploiting the Many Sides of Target Row RefreshPietro Frigo, Emanuele Vannacci, Hasan Hassan, Victor van der Veen et al.S&P 2020 · 274 citations
- BLACKSMITH: Scalable Rowhammering in the Frequency DomainPatrick Jattke, Victor van der Veen, Pietro Frigo, Stijn Gunter et al.S&P 2022 · 140 citations
- ProTRR: Principled yet Optimal In-DRAM Target Row RefreshMichele Marazzi, Patrick Jattke, Flavien Solt, Kaveh RazaviS&P 2022 · 101 citations
