Characterizing and Mitigating Soft Errors in GPU DRAM
Michael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee, Paul Racunas, Saurabh Hukerikar, Timothy Tsai, Siva Kumar Sastry Hari, Stephen W. Keckler
摘要
GPUs are used in high-reliability systems, including high-performance computers and autonomous vehicles. Because GPUs employ a high-bandwidth, wide-interface to DRAM and fetch each memory access from a single DRAM device, implementing full-device correction through ECC is expensive and impractical. This challenge is compounded by worsening relative rates of multi-bit DRAM errors and increasing GPU memory capacities. This paper first presents high-energy neutron beam testing results for the HBM2 memory on a compute-class GPU. These results uncovered unexpected intermittent errors that we determine to be caused by cell damage from the high-intensity beam. As these errors are an artifact of the testing apparatus, we provide best-practice guidance on how to identify and filter them from the results of beam testing campaigns. Second, we use the soft error beam testing results to inform the design and evaluation of system-level error protection mechanisms by reporting the relative error rates and error patterns from soft errors in GPU DRAM. We observe locality in the multi-bit errors, which we attribute to the underlying structure of the HBM2 memory. Based on these error patterns, we propose several novel ECC schemes to decrease the silent data corruption risk by up to five orders of magnitude relative to SEC-DED ECC, while also reducing the number of uncorrectable errors by up to 7.87 ×. We compare novel binary and symbol-based ECC organizations that differ in their design complexity, hardware overheads, and permanent error correction abilities, ultimately recommending two promising organizations. These schemes replace SEC-DED ECC with no additional redundancy, likely no performance impacts, and modest area and complexity costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- RowPress: Amplifying Read Disturbance in Modern DRAM ChipsHaocong Luo, Ataberk Olgun, Abdullah Giray Yaglikçi, Yahya Can Tugrul 等ISCA 2023 · 被引用 71 次
- Unity ECC: Unified Memory Protection Against Bit and Chip ErrorsDongwhee Kim, Jaeyoon Lee, Wonyeong Jung, Michael B. Sullivan 等SC 2023 · 被引用 23 次
- Understanding the Effects of Permanent Faults in GPU's Parallelism Management and Control UnitsJuan-David Guerrero-Balaguera, Josie Esteban Rodriguez Condia, Fernando Fernandes dos Santos, Matteo Sonza Reorda 等SC 2023 · 被引用 18 次
- Implicit Memory Tagging: No-Overhead Memory Safety Using Alias-Free Tagged ECCMichael B. Sullivan, Mohamed Tarek Ibn Ziad, Aamer Jaleel, Stephen W. KecklerISCA 2023 · 被引用 17 次
- Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers ModelsLucas Roquet, Fernando Fernandes dos Santos, Paolo Rech, Marcello Traiola 等DAC 2024 · 被引用 15 次
相关 Paper
- PoP-ECC: Robust and Flexible Error Correction against Multi-Bit Upsets in DNN AcceleratorsTaewon Park, Saeid Gorgin, Dongwhee Kim, Jaeho Shin 等DAC 2025 · 被引用 2 次
- Structural Coding: A Low-Cost Scheme to Protect CNNs from Large-Granularity Memory FaultsAli Asgari Khoshouyeh, Florian Geissler, Syed Sha Qutub, Michael Paulitsch 等SC 2023 · 被引用 8 次
- HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECCRuizhi Zhu, Yanan Guo, Huize Li, Weidong Cao 等ISCA 2026
- Compiler-directed soft error resilience for lightweight GPU register file protectionHongjune Kim, Jianping Zeng, Qingrui Liu, Mohammad Abdel-Majeed 等PLDI 2020 · 被引用 31 次
- ECC Enabled Reliable and Performant Processing-in-MemoryJeageun Jung, Margaret Lee, Mattan ErezISCA 2026
