RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
Hanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae Kim
摘要
As DRAM scales in density and adopts 3D integration, raw fault rates increase and multi-bit errors are no longer rare. Such errors can severely impact Deep Neural Networks (DNNs): although DNNs tolerate small numerical perturbations, random bit flips can create extreme outliers that propagate and sharply degrade accuracy. Large Language Models (LLMs) are particularly vulnerable because attention, residual, and normalization layers can amplify and preserve a single corrupted activation across many layers, destabilizing inference.
This paper introduces RangeGuard, a metadata-centric errorcorrecting framework that provides strong reliability and high efficiency based on bounded approximate correction. Instead of protecting raw bits, RangeGuard encodes compact Range Identifiers (RIDs) that capture the numerical range of each value. These compact metadata enable efficient use of limited redundancy and concentrate protection on range changes-which indicate harmful semantic deviations-while ignoring benign intra-range variations. Upon detecting a range change, RangeGuard restores the correct range and substitutes a representative value, ensuring that error magnitudes are bounded within the range. Based on RIDs, RangeGuard can tolerate 64+ bits of error using only 16 bits of parity available in GPU memories without a noticeable accuracy loss. By introducing semantic range protection, RangeGuard enables reliable DNN execution even under frequent memory errors and tight redundancy budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault AttacksSanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida 等USENIX Security 2019 · 被引用 255 次
- FIdelity: Efficient Resilience Analysis Framework for Deep Learning AcceleratorsYi He, Prasanna Balaprakash, Yanjing LiMICRO 2020 · 被引用 82 次
- Characterizing and Mitigating Soft Errors in GPU DRAMMichael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee 等MICRO 2021 · 被引用 45 次
- Unity ECC: Unified Memory Protection Against Bit and Chip ErrorsDongwhee Kim, Jaeyoon Lee, Wonyeong Jung, Michael B. Sullivan 等SC 2023 · 被引用 23 次
相关 Paper
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu 等EMNLP 2025
- Structural Coding: A Low-Cost Scheme to Protect CNNs from Large-Granularity Memory FaultsAli Asgari Khoshouyeh, Florian Geissler, Syed Sha Qutub, Michael Paulitsch 等SC 2023 · 被引用 8 次
- BitShield: Defending Against Bit-Flip Attacks on DNN ExecutablesYanzuo Chen, Yuanyuan Yuan, Zhibo Liu, Sihang Hu 等NDSS 2025
- SafeGuard: Reducing the Security Risk from Row-Hammer via Low-Cost Integrity ProtectionAli Fakhrzadehgan, Yale N. Patt, Prashant J. Nair, Moinuddin K. QureshiHPCA 2022 · 被引用 51 次
- SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeQinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang 等ASE 2025 · 被引用 1 次
