SAVE: Software-Implemented Fault Tolerance for Model Inference against GPU Memory Bit Flips
Wenxin Zheng, Bin Xu, Jinyu Gu, Haibo Chen
Abstract
Machine learning models are used in safety-critical edge applications such as autonomous driving, industrial robots, and satellites. However, GPU memory bit flips can significantly reduce the model accuracy. Existing mitigations either compromise accuracy or introduce substantial overhead.
Our insight is that not all hardware bits are created equal and bit flips vary in their impact on model inference. Specifically, for the GPU memory, modern AI accelerators provide bit-flip-free but small reliable memory. For the model inference, due to nonlinear activation functions in the model, some bits are naturally robust against flips, while other vulnerable bits can silently corrupt results. Thus, we prioritize the allocation of vulnerable bits' computations in the reliable memory to enhance the robustness of the model inference.
We propose SAVE, a software-implemented fault tolerance system that protects model inference without modifying the model and with minimal performance impact. SAVE operates in four stages: Selection to identify vulnerable bits based on the intrinsic characteristics of model inference, Allocation to prioritize computations related to more vulnerable bits in reliable memory, Verification to efficiently detect errors through asynchronous CPU checks, and Edit to recover from detected faults. Evaluation across computer vision, robotics, and decision-making models shows that SAVE maintains model accuracy even under 4K bit flips while incurring less than 9% performance overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a736ee89-9237-46be-82e6-bc70ee20c8d0Cited by top-tier papers1
Ask how each one uses itBuilds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Drammer: Deterministic Rowhammer Attacks on Mobile PlatformsVictor van der Veen, Yanick Fratantonio, Martina Lindorfer, Daniel Gruss et al.CCS 2016 · 381 citations
- Bit-Flip Attack: Crushing Neural Network With Progressive Bit SearchAdnan Siraj Rakin, Zhezhi He, Deliang FanICCV 2019 · 309 citations
Related papers
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault AttacksSanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida et al.USENIX Security 2019 · 255 citations
- Arithmetic-intensity-guided fault tolerance for neural network inference on GPUsJack Kosaian, K. V. RashmiSC 2021 · 51 citations
- Pruning of Deep Neural Networks for Fault-Tolerant Memristor-based AcceleratorsChing-Yuan Chen, Krishnendu ChakrabartyDAC 2021 · 24 citations
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu et al.EMNLP 2025
- Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers ModelsLucas Roquet, Fernando Fernandes dos Santos, Paolo Rech, Marcello Traiola et al.DAC 2024 · 15 citations
