AI-Guardian: Defeating Adversarial Attacks using Backdoors
Hong Zhu, Shengzhi Zhang, Kai Chen
Abstract
Deep neural networks (DNNs) have been widely used in many fields due to their increasingly high accuracy. However, they are also vulnerable to adversarial attacks, posing a serious threat to security-critical applications such as autonomous driving, remote diagnosis, etc. Existing solutions are limited in detecting/preventing such attacks, and also impacting the performance on the original tasks. In this paper, we present AI-Guardian, a novel approach to defeating adversarial attacks that leverages intentionally embedded backdoors to fail the adversarial perturbations and maintain the performance of the original main task. We extensively evaluate AI-Guardian using five popular adversarial example generation approaches, and experimental results demonstrate its efficacy in defeating adversarial attacks. Specifically, AI-Guardian reduces the attack success rate from 97.3% to 3.2%, which outperforms the state-of-the-art works by 30.9%, with only a 0.9% decline on the clean data accuracy. Furthermore, AI-Guardian introduces only 0.36% overhead to the model prediction time, almost negligible in most cases.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6272b3df-4633-4b57-a4d7-9b1059534294Cited by top-tier papers5
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 20 citations
- PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency ReasoningZhiyuan Yu, Ao Li, Ruoyao Wen, Yijia Chen et al.CCS 2024 · 4 citations
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- Robust Adversarial Attacks Against Unknown Disturbance via Inverse Gradient SampleZhaoyang Zhang, Shen Wang, Runze Liu, Guopu Zhu et al.ICLR 2026
- Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property RefinementJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangCCS 2025
Related papers
- AdvDoor: adversarial backdoor attack of deep learning systemQuan Zhang, Yifeng Ding, Yongqiang Tian, Jianmin Guo et al.ISSTA 2021 · 57 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- PatchBackdoor: Backdoor Attack against Deep Neural Networks without Model ModificationYizhen Yuan, Rui Kong, Shenghao Xie, Yuanchun Li et al.ACM MM 2023 · 12 citations
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
- Adversarial-Inspired Backdoor Defense via Bridging Backdoor and Adversarial AttacksJia-Li Yin, Weijian Wang, Lyhwa, Wei Lin et al.AAAI 2025 · 9 citations
