Trap and Replace: Defending Backdoor Attacks by Trapping Them into an Easy-to-Replace Subnetwork
Haotao Wang, Junyuan Hong, Aston Zhang, Jiayu Zhou, Zhangyang Wang
摘要
Deep neural networks (DNNs) are vulnerable to backdoor attacks. Previous works have shown it extremely challenging to unlearn the undesired backdoor behavior from the network, since the entire network can be affected by the backdoor samples. In this paper, we propose a brand-new backdoor defense strategy, which makes it much easier to remove the harmful influence of backdoor samples from the model. Our defense strategy, Trap and Replace, consists of two stages. In the first stage, we bait and trap the backdoors in a small and easy-to-replace subnetwork. Specifically, we add an auxiliary image reconstruction head on top of the stem network shared with a light-weighted classification head. The intuition is that the auxiliary image reconstruction task encourages the stem network to keep sufficient low-level visual features that are hard to learn but semantically correct, instead of overfitting to the easy-to-learn but semantically incorrect backdoor correlations. As a result, when trained on backdoored datasets, the backdoors are easily baited towards the unprotected classification head, since it is much more vulnerable than the shared stem, leaving the stem network hardly poisoned. In the second stage, we replace the poisoned light-weighted classification head with an untainted one, by re-training it from scratch only on a small holdout dataset with clean samples, while fixing the stem network. As a result, both the stem and the classification head in the final network are hardly affected by backdoor training samples. We evaluate our method against ten different backdoor attacks. Our method outperforms previous state-of-the-art methods by up to 20.57%, 9.80%, and 13.72% attack success rate and on-average 3.14%, 1.80%, and 1.21% clean classification accuracy on CIFAR10, GTSRB, and ImageNet-12, respectively. Code is available at https: //github.com/VITA-Group/Trap-and-Replace-Backdoor-Defense .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan 等NeurIPS 2023 · 被引用 66 次
- Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through HoneypotsRuixiang (Ryan) Tang, Jiayi Yuan, Yiming Li, Zirui Liu 等NeurIPS 2023 · 被引用 31 次
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu 等ICML 2023 · 被引用 16 次
- On the Difficulty of Defending Contrastive Learning against Backdoor AttacksChangjiang Li, Ren Pang, Bochuan Cao, Zhaohan Xi 等USENIX Security 2024 · 被引用 10 次
- Backdoor Defense via Enhanced Splitting and Trap IsolationHongrui Yu, Lu Qi, Wanyu Lin, Jian Chen 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper30
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
相关 Paper
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
- Need for Speed: Taming Backdoor Attacks with Speed and PrecisionZhuo Ma, Yilong Yang, Yang Liu, Tong Yang 等S&P 2024 · 被引用 6 次
- Backdoor Defense via Deconfounded Representation LearningZaixi Zhang, Qi Liu, Zhicai Wang, Zepu Lu 等CVPR 2023
- Neural Polarizer: A Lightweight and Effective Backdoor Defense via Purifying Poisoned FeaturesMingli Zhu, Shaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 被引用 68 次
- BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input DetectionTinghao Xie, Xiangyu Qi, Ping He, Yiming Li 等ICLR 2024 · 被引用 20 次
