RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
Hanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae Kim
Abstract
As DRAM scales in density and adopts 3D integration, raw fault rates increase and multi-bit errors are no longer rare. Such errors can severely impact Deep Neural Networks (DNNs): although DNNs tolerate small numerical perturbations, random bit flips can create extreme outliers that propagate and sharply degrade accuracy. Large Language Models (LLMs) are particularly vulnerable because attention, residual, and normalization layers can amplify and preserve a single corrupted activation across many layers, destabilizing inference.
This paper introduces RangeGuard, a metadata-centric errorcorrecting framework that provides strong reliability and high efficiency based on bounded approximate correction. Instead of protecting raw bits, RangeGuard encodes compact Range Identifiers (RIDs) that capture the numerical range of each value. These compact metadata enable efficient use of limited redundancy and concentrate protection on range changes-which indicate harmful semantic deviations-while ignoring benign intra-range variations. Upon detecting a range change, RangeGuard restores the correct range and substitutes a representative value, ensuring that error magnitudes are bounded within the range. Based on RIDs, RangeGuard can tolerate 64+ bits of error using only 16 bits of parity available in GPU memories without a noticeable accuracy loss. By introducing semantic range protection, RangeGuard enables reliable DNN execution even under frequent memory errors and tight redundancy budgets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3b046b1-1e32-426b-bbf3-023b20ac77a8Builds on12
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault AttacksSanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida et al.USENIX Security 2019 · 255 citations
- FIdelity: Efficient Resilience Analysis Framework for Deep Learning AcceleratorsYi He, Prasanna Balaprakash, Yanjing LiMICRO 2020 · 82 citations
- Characterizing and Mitigating Soft Errors in GPU DRAMMichael B. Sullivan, Nirmal R. Saxena, Mike O'Connor, Donghyuk Lee et al.MICRO 2021 · 45 citations
- Unity ECC: Unified Memory Protection Against Bit and Chip ErrorsDongwhee Kim, Jaeyoon Lee, Wonyeong Jung, Michael B. Sullivan et al.SC 2023 · 23 citations
Related papers
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu et al.EMNLP 2025
- Structural Coding: A Low-Cost Scheme to Protect CNNs from Large-Granularity Memory FaultsAli Asgari Khoshouyeh, Florian Geissler, Syed Sha Qutub, Michael Paulitsch et al.SC 2023 · 8 citations
- BitShield: Defending Against Bit-Flip Attacks on DNN ExecutablesYanzuo Chen, Yuanyuan Yuan, Zhibo Liu, Sihang Hu et al.NDSS 2025
- SafeGuard: Reducing the Security Risk from Row-Hammer via Low-Cost Integrity ProtectionAli Fakhrzadehgan, Yale N. Patt, Prashant J. Nair, Moinuddin K. QureshiHPCA 2022 · 51 citations
- SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeQinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang et al.ASE 2025 · 1 citation
