Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
Rui Min, Zeyu Qin, Nevin L. Zhang, Li Shen, Minhao Cheng
摘要
Backdoor attacks pose a significant threat to Deep Neural Networks (DNNs) as they allow attackers to manipulate model predictions with backdoor triggers. To address these security vulnerabilities, various backdoor purification methods have been proposed to purify compromised models. Typically, these purified models exhibit low Attack Success Rates (ASR), rendering them resistant to backdoored inputs. However, Does achieving a low ASR through current safety purification methods truly eliminate learned backdoor features from the pretraining phase? In this paper, we provide an affirmative answer to this question by thoroughly investigating the Post-Purification Robustness of current backdoor purification methods. We find that current safety purification methods are vulnerable to the rapid re-learning of backdoor behavior, even when further fine-tuning of purified models is performed using a very small number of poisoned samples. Based on this, we further propose the practical Query-based Reactivation Attack (QRA) which could effectively reactivate the backdoor by merely querying purified models. We find the failure to achieve satisfactory post-purification robustness stems from the insufficient deviation of purified models from the backdoored model along the backdoor-connected path. To improve the post-purification robustness, we propose a straightforward tuning defense, Path-Aware Minimization (PAM), which promotes deviation along backdoor-connected paths with extra model updates. Extensive experiments demonstrate that PAM significantly improves post-purification robustness while maintaining a good clean accuracy and low ASR. Our work provides a new perspective on understanding the effectiveness of backdoor safety tuning and highlights the importance of faithfully assessing the model's safety.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion ModelsYingsha Xie, Rui Min, Zeyu Qin, Fei Ma 等NeurIPS 2025 · 被引用 3 次
- From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model MergingZhenqian Zhu, Yamin Hu, Yiya Diao, Weixiang Li 等ICML 2026 · 被引用 1 次
- Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful PerturbationTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin 等ICLR 2025
- Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual LearningZhenqian Zhu, Yamin Hu, Yujiang Liu, Luping Wei 等CCS 2026
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
相关 Paper
- Towards Stable Backdoor Purification through Feature Shift TuningRui Min, Zeyu Qin, Li Shen, Minhao ChengNeurIPS 2023 · 被引用 43 次
- From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural NetworksXianda Zhang, Baolin Zheng, Jianbao Hu, Chengyang Li 等AAAI 2024 · 被引用 3 次
- Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor ActivenessWeilin Lin, Li Liu, Shaokui Wei, Jianze Li 等NeurIPS 2024 · 被引用 16 次
- Redeem Myself: Purifying Backdoors in Deep Learning Models using Self Attention DistillationXueluan Gong, Yanjiao Chen, Wang Yang, Qian Wang 等S&P 2023
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
