USENIX ATC2025顶会
SAVE: Software-Implemented Fault Tolerance for Model Inference against GPU Memory Bit Flips
Wenxin Zheng, Bin Xu, Jinyu Gu, Haibo Chen
摘要
Machine learning models are used in safety-critical edge applications such as autonomous driving, industrial robots, and satellites. However, GPU memory bit flips can significantly reduce the model accuracy. Existing mitigations either compromise accuracy or introduce substantial overhead.
Our insight is that not all hardware bits are created equal and bit flips vary in their impact on model inference. Specifically, for the GPU memory, modern AI accelerators provide bit-flip-free but small reliable memory. For the model inference, due to nonlinear activation functions in the model, some bits are naturally robust against flips, while other vulnerable bits can silently corrupt results. Thus, we prioritize the allocation of vulnerable bits' computations in the reliable memory to enhance the robustness of the model inference.
We propose SAVE, a software-implemented fault tolerance system that protects model inference without modifying the model and with minimal performance impact. SAVE operates in four stages: Selection to identify vulnerable bits based on the intrinsic characteristics of model inference, Allocation to prioritize computations related to more vulnerable bits in reliable memory, Verification to efficiently detect errors through asynchronous CPU checks, and Edit to recover from detected faults. Evaluation across computer vision, robotics, and decision-making models shows that SAVE maintains model accuracy even under 4K bit flips while incurring less than 9% performance overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Drammer: Deterministic Rowhammer Attacks on Mobile PlatformsVictor van der Veen, Yanick Fratantonio, Martina Lindorfer, Daniel Gruss 等CCS 2016 · 被引用 381 次
- Bit-Flip Attack: Crushing Neural Network With Progressive Bit SearchAdnan Siraj Rakin, Zhezhi He, Deliang FanICCV 2019 · 被引用 309 次
相关 Paper
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault AttacksSanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida 等USENIX Security 2019 · 被引用 255 次
- Arithmetic-intensity-guided fault tolerance for neural network inference on GPUsJack Kosaian, K. V. RashmiSC 2021 · 被引用 51 次
- Pruning of Deep Neural Networks for Fault-Tolerant Memristor-based AcceleratorsChing-Yuan Chen, Krishnendu ChakrabartyDAC 2021 · 被引用 24 次
- Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu 等EMNLP 2025
- Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers ModelsLucas Roquet, Fernando Fernandes dos Santos, Paolo Rech, Marcello Traiola 等DAC 2024 · 被引用 15 次
