Characterizing and Mitigating Soft Errors in GPU DRAM
Michael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee, Paul Racunas, Saurabh Hukerikar, Timothy Tsai, Siva Kumar Sastry Hari, Stephen W. Keckler
Abstract
GPUs are used in high-reliability systems, including high-performance computers and autonomous vehicles. Because GPUs employ a high-bandwidth, wide-interface to DRAM and fetch each memory access from a single DRAM device, implementing full-device correction through ECC is expensive and impractical. This challenge is compounded by worsening relative rates of multi-bit DRAM errors and increasing GPU memory capacities. This paper first presents high-energy neutron beam testing results for the HBM2 memory on a compute-class GPU. These results uncovered unexpected intermittent errors that we determine to be caused by cell damage from the high-intensity beam. As these errors are an artifact of the testing apparatus, we provide best-practice guidance on how to identify and filter them from the results of beam testing campaigns. Second, we use the soft error beam testing results to inform the design and evaluation of system-level error protection mechanisms by reporting the relative error rates and error patterns from soft errors in GPU DRAM. We observe locality in the multi-bit errors, which we attribute to the underlying structure of the HBM2 memory. Based on these error patterns, we propose several novel ECC schemes to decrease the silent data corruption risk by up to five orders of magnitude relative to SEC-DED ECC, while also reducing the number of uncorrectable errors by up to 7.87 ×. We compare novel binary and symbol-based ECC organizations that differ in their design complexity, hardware overheads, and permanent error correction abilities, ultimately recommending two promising organizations. These schemes replace SEC-DED ECC with no additional redundancy, likely no performance impacts, and modest area and complexity costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3edb73cd-3483-4ff0-84e5-d10f6321f3a7Cited by top-tier papers8
- RowPress: Amplifying Read Disturbance in Modern DRAM ChipsHaocong Luo, Ataberk Olgun, Abdullah Giray Yaglikçi, Yahya Can Tugrul et al.ISCA 2023 · 71 citations
- Unity ECC: Unified Memory Protection Against Bit and Chip ErrorsDongwhee Kim, Jaeyoon Lee, Wonyeong Jung, Michael B. Sullivan et al.SC 2023 · 23 citations
- Understanding the Effects of Permanent Faults in GPU's Parallelism Management and Control UnitsJuan-David Guerrero-Balaguera, Josie Esteban Rodriguez Condia, Fernando Fernandes dos Santos, Matteo Sonza Reorda et al.SC 2023 · 18 citations
- Implicit Memory Tagging: No-Overhead Memory Safety Using Alias-Free Tagged ECCMichael B. Sullivan, Mohamed Tarek Ibn Ziad, Aamer Jaleel, Stephen W. KecklerISCA 2023 · 17 citations
- Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers ModelsLucas Roquet, Fernando Fernandes dos Santos, Paolo Rech, Marcello Traiola et al.DAC 2024 · 15 citations
Related papers
- PoP-ECC: Robust and Flexible Error Correction against Multi-Bit Upsets in DNN AcceleratorsTaewon Park, Saeid Gorgin, Dongwhee Kim, Jaeho Shin et al.DAC 2025 · 2 citations
- Structural Coding: A Low-Cost Scheme to Protect CNNs from Large-Granularity Memory FaultsAli Asgari Khoshouyeh, Florian Geissler, Syed Sha Qutub, Michael Paulitsch et al.SC 2023 · 8 citations
- HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECCRuizhi Zhu, Yanan Guo, Huize Li, Weidong Cao et al.ISCA 2026
- Compiler-directed soft error resilience for lightweight GPU register file protectionHongjune Kim, Jianping Zeng, Qingrui Liu, Mohammad Abdel-Majeed et al.PLDI 2020 · 31 citations
- ECC Enabled Reliable and Performant Processing-in-MemoryJeageun Jung, Margaret Lee, Mattan ErezISCA 2026
