Revisiting the Assumption of Latent Separability for Backdoor Defenses
Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, Prateek Mittal
摘要
Recent studies revealed that deep learning is susceptible to backdoor poisoning attacks. An adversary can embed a hidden backdoor into a model to manipulate its predictions by only modifying a few training data, without controlling the training process. Currently, a tangible signature has been widely observed across a diverse set of backdoor poisoning attacks -models trained on a poisoned dataset tend to learn separable latent representations for poison and clean samples. This latent separation is so pervasive that a family of backdoor defenses directly take it as a default assumption (dubbed latent separability assumption), based on which to identify poison samples via cluster analysis in the latent space. An intriguing question consequently follows: is the latent separation unavoidable for backdoor poisoning attacks? This question is central to understanding whether the assumption of latent separability provides a reliable foundation for defending against backdoor poisoning attacks. In this paper, we design adaptive backdoor poisoning attacks to present counter-examples against this assumption. Our methods include two key components: (1) a set of triggerplanted samples correctly labeled to their semantic classes (other than the target class) that can regularize backdoor learning; (2) asymmetric trigger planting strategies that help to boost attack success rate (ASR) as well as to diversify latent representations of poison samples. Extensive experiments on benchmark datasets verify the effectiveness of our adaptive attacks in bypassing existing latent separation based defenses. Our codes are available at https: //github.com/Unispac/Circumventing-Backdoor-Defenses .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia 等NeurIPS 2023 · 被引用 93 次
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 被引用 69 次
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan 等NeurIPS 2023 · 被引用 66 次
它引用的顶会 Paper17
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- DBA: Distributed Backdoor Attacks against Federated LearningChulin Xie, Keli Huang, Pin-Yu Chen, Bo LiICLR 2020 · 被引用 901 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
相关 Paper
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong 等CVPR 2022 · 被引用 72 次
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu 等CCS 2023 · 被引用 170 次
- Adversarial-Inspired Backdoor Defense via Bridging Backdoor and Adversarial AttacksJia-Li Yin, Weijian Wang, Lyhwa, Wei Lin 等AAAI 2025 · 被引用 9 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
