Improving Robustness Against Stealthy Weight Bit-Flip Attacks by Output Code Matching
Ozan Özdenizci, Robert Legenstein
Abstract
Deep neural networks (DNNs) have been shown to be vulnerable against adversarial weight bit-flip attacks through hardware-induced fault-injection methods on the memory systems where network parameters are stored. Recent attacks pose the further concerning threat of finding minimal targeted and stealthy weight bit-flips that preserve expected behavior for untargeted test samples. This renders the attack undetectable from a DNN operation perspective. We propose a DNN defense mechanism to improve robustness in such realistic stealthy weight bit-flip attack scenarios. Our output code matching networks use an output coding scheme where the usual one-hot encoding of classes is replaced by partially overlapping bit strings. We show that this encoding significantly reduces attack stealthiness. Importantly, our approach is compatible with existing defenses and DNN architectures. It can be efficiently implemented on pre-trained models by simply re-defining the output classification layer and finetuning. Experimental benchmark evaluations show that output code matching is superior to existing regularized weight quantization based defenses, and an effective defense against stealthy weight bit-flip attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d87f113-ee3d-4d7c-a65c-fa5e48253315Cited by top-tier papers3
- Crossfire: An Elastic Defense Framework for Graph Neural Networks Under Bit Flip AttacksLorenz Kummer, Samir Moustafa, Wilfried N. Gansterer, Nils Morten KriegeAAAI 2025
- Diversity-aware Weight Perturbation Promotes Robust AdaptationZibo Chen, Ruxin Li, Zilu WangICML 2026
- Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor ManipulationJialai Wang, Yuxiao Wu, Weiye Xu, Yating Huang et al.CVPR 2025
Builds on17
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Drammer: Deterministic Rowhammer Attacks on Mobile PlatformsVictor van der Veen, Yanick Fratantonio, Martina Lindorfer, Daniel Gruss et al.CCS 2016 · 381 citations
- Bit-Flip Attack: Crushing Neural Network With Progressive Bit SearchAdnan Siraj Rakin, Zhezhi He, Deliang FanICCV 2019 · 309 citations
- Another Flip in the Wall of Rowhammer DefensesDaniel Gruss, Moritz Lipp, Michael Schwarz, Daniel Genkin et al.S&P 2018 · 288 citations
Related papers
- Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN ExecutablesYanzuo Chen, Zhibo Liu, Yuanyuan Yuan, Sihang Hu et al.NDSS 2025
- Defending and Harnessing the Bit-Flip Based Adversarial Weight AttackZhezhi He, Adnan Siraj Rakin, Jingtao Li, Chaitali Chakrabarti et al.CVPR 2020
- One-bit Flip is All You Need: When Bit-flip Attack Meets Model TrainingJianshuo Dong, Han Qiu, Yiming Li, Tianwei Zhang et al.ICCV 2023 · 33 citations
- Defending Bit-Flip Attack through DNN Weight ReconstructionJingtao Li, Adnan Siraj Rakin, Yan Xiong, Liangliang Chang et al.DAC 2020 · 55 citations
- DNN-Defender: A Victim-Focused In-DRAM Defense Mechanism for Taming Adversarial Weight Attack on DNNsRanyang Zhou, Sabbir Ahmed, Adnan Siraj Rakin, Shaahin AngiziDAC 2024 · 2 citations
