Marksman Backdoor: Backdoor Attacks with Arbitrary Target Class
Khoa D. Doan, Yingjie Lao, Ping Li
摘要
In recent years, machine learning models have been shown to be vulnerable to backdoor attacks. Under such attacks, an adversary embeds a stealthy backdoor into the trained model such that the compromised models will behave normally on clean inputs but will misclassify according to the adversary's control on maliciously constructed input with a trigger. While these existing attacks are very effective, the adversary's capability is limited: given an input, these attacks can only cause the model to misclassify toward a single pre-defined or target class. In contrast, this paper exploits a novel backdoor attack with a much more powerful payload, denoted as Marksman, where the adversary can arbitrarily choose which target class the model will misclassify given any input during inference. To achieve this goal, we propose to represent the trigger function as a class-conditional generative model and to inject the backdoor in a constrained optimization framework, where the trigger function learns to generate an optimal trigger pattern to attack any target class at will while simultaneously embedding this generative backdoor into the trained model. Given the learned trigger-generation function, during inference, the adversary can specify an arbitrary backdoor attack target class, and an appropriate trigger causing the model to classify toward this target class is created accordingly. We show empirically that the proposed framework achieves high attack performance (e.g., 100% attack success rates in several experiments) while preserving the cleandata performance in several benchmark datasets, including MNIST, CIFAR10, GTSRB, and TinyImageNet. The proposed Marksman backdoor attack can also easily bypass existing backdoor defenses that were originally designed against backdoor attacks with a single target class. Our work takes another significant step toward understanding the extensive risks of backdoor attacks in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- IBA: Towards Irreversible Backdoor Attacks in Federated LearningThuy Dung Nguyen, Tuan Nguyen, Anh Tran, Khoa D. Doan 等NeurIPS 2023 · 被引用 94 次
- Defending Backdoor Attacks on Vision Transformer via Patch ProcessingKhoa D. Doan, Yingjie Lao, Peng Yang, Ping LiAAAI 2023 · 被引用 33 次
- IAG: Input-aware Backdoor Attack on VLM-based Visual GroundingJunxian Li, Beining Xu, Simin Chen, Jiatong Li 等CVPR 2026 · 被引用 13 次
- Data Free Backdoor AttacksBochuan Cao, Jinyuan Jia, Chuxuan Hu, Wenbo Guo 等NeurIPS 2024 · 被引用 12 次
- BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial ManipulationRui Chu, Bingyin Zhao, Hanling Jiang, Shuchin Aeron 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper19
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
相关 Paper
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu 等CCS 2023 · 被引用 170 次
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong 等CVPR 2022 · 被引用 72 次
- LIRA: Learnable, Imperceptible and Robust Backdoor AttacksKhoa D. Doan, Yingjie Lao, Weijie Zhao, Ping LiICCV 2021 · 被引用 313 次
- COMBAT: Alternated Training for Effective Clean-Label Backdoor AttacksTran Huynh, Dang Nguyen, Tung Pham, Anh TranAAAI 2024 · 被引用 25 次
- Single Image Backdoor Inversion via Robust Smoothed ClassifiersMingjie Sun, Zico KolterCVPR 2023
