Few-shot Backdoor Defense Using Shapley Estimation
Jiyang Guan, Zhuozhuo Tu, Ran He, Dacheng Tao
Abstract
Deep neural networks have achieved impressive performance in a variety of tasks over the last decade, such as autonomous driving, face recognition, and medical diagnosis. However, prior works show that deep neural networks are easily manipulated into specific, attacker-decided behaviors in the inference stage by backdoor attacks which inject malicious small hidden triggers into model training, raising serious security threats. To determine the triggered neurons and protect against backdoor attacks, we exploit Shapley value and develop a new approach called Shapley Pruning (ShapPruning) that successfully mitigates backdoor attacks from models in a data-insufficient situation (1 image per class or even free of data). Considering the interaction between neurons, ShapPruning identifies the few infected neurons (under 1 % of all neurons) and manages to protect the model's structure and accuracy after pruning as many infected neurons as possible. To accelerate ShapPruning, we further propose discarding threshold and ∊ -greedy strategy to accelerate Shapley estimation, making it possible to repair poisoned models with only several minutes. Experiments demonstrate the effectiveness and robustness of our method against various attacks and tasks compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Reconstructive Neuron Pruning for Backdoor DefenseYige Li, Xixiang Lyu, Xingjun Ma, Nodens Koren et al.ICML 2023 · 86 citations
- Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural NetworksJiyang Guan, Jian Liang, Ran HeNeurIPS 2022 · 57 citations
- UMD: Unsupervised Model Detection for X2X Backdoor AttacksZhen Xiang, Zidi Xiong, Bo LiICML 2023 · 27 citations
- Does Few-Shot Learning Suffer from Backdoor Attacks?Xinwei Liu, Xiaojun Jia, Jindong Gu, Yuan Xun et al.AAAI 2024 · 24 citations
- Inspecting Prediction Confidence for Detecting Black-Box Backdoor AttacksTong Wang, Yuan Yao, Feng Xu, Miao Xu et al.AAAI 2024 · 15 citations
Builds on15
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson et al.AAAI 2020 · 210 citations
Related papers
- Backdoor Defense via Test-Time Detecting and RepairingJiyang Guan, Jian Liang, Ran HeCVPR 2024
- Adversarial Neuron Pruning Purifies Backdoored Deep ModelsDongxian Wu, Yisen WangNeurIPS 2021 · 441 citations
- Adversarial Feature Map Pruning for BackdoorDong Huang, Qingwen BuICLR 2024 · 6 citations
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang et al.ICCV 2021 · 128 citations
- Need for Speed: Taming Backdoor Attacks with Speed and PrecisionZhuo Ma, Yilong Yang, Yang Liu, Tong Yang et al.S&P 2024 · 6 citations
