Breaking the False Sense of Security in Backdoor Defense through Re-Activation Attack
Mingli Zhu, Siyuan Liang, Baoyuan Wu
摘要
Deep neural networks face persistent challenges in defending against backdoor attacks, leading to an ongoing battle between attacks and defenses. While existing backdoor defense strategies have shown promising performance on reducing attack success rates, can we confidently claim that the backdoor threat has truly been eliminated from the model? To address it, we re-investigate the characteristics of the backdoored models after defense (denoted as defense models). Surprisingly, we find that the original backdoors still exist in defense models derived from existing post-training defense strategies, and the backdoor existence is measured by a novel metric called backdoor existence coefficient. It implies that the backdoors just lie dormant rather than being eliminated. To further verify this finding, we empirically show that these dormant backdoors can be easily re-activated during inference, by manipulating the original trigger with well-designed tiny perturbation using universal adversarial attack. More practically, we extend our backdoor reactivation to black-box scenario, where the defense model can only be queried by the adversary during inference, and develop two effective methods, i.e., query-based and transfer-based backdoor re-activation attacks. The effectiveness of the proposed methods are verified on both image classification and multimodal contrastive learning (i.e., CLIP) tasks. In conclusion, this work uncovers a critical vulnerability that has never been explored in existing defense strategies, emphasizing the urgency of designing more robust and advanced backdoor defense mechanisms in the future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 被引用 20 次
- ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language ModelsXuxu Liu, Siyuan Liang, Mengya Han, Yong Luo 等ACL 2025 · 被引用 13 次
- Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive LearningXinwei Liu, Xiaojun Jia, Yuan Xun, Siyuan Liang 等ACM MM 2024 · 被引用 11 次
- Lie Detector: Unified Backdoor Detection via Cross-Examination FrameworkXuan Wang, Siyuan Liang, Dongping Liao, Han Fang 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
相关 Paper
- REFINE: Inversion-Free Backdoor Defense via Model ReprogrammingYukun Chen, Shuo Shao, Enhao Huang, Yiming Li 等ICLR 2025
- Test-Time Multimodal Backdoor Detection by Contrastive PromptingYuwei Niu, Shuo He, Qi Wei, Zongyu Wu 等ICML 2025
- BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive LearningSiyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu 等CVPR 2024
- DeDe: Detecting Backdoor Samples for SSL Encoders via DecodersSizai Hou, Songze Li, Duanyi YaoCVPR 2025
- CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset SeparationBinyan Xu, Fan Yang, Xilin Dai, Di Tang 等ACM MM 2025 · 被引用 12 次
