Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
Shawn Shan, Emily Wenger, Bolun Wang, Bo Li, Haitao Zheng, Ben Y. Zhao
摘要
Deep neural networks (DNN) are known to be vulnerable to adversarial attacks. Numerous efforts either try to patch weaknesses in trained models, or try to make it difficult or costly to compute adversarial examples that exploit them. In our work, we explore a new "honeypot" approach to protect DNN models. We intentionally inject trapdoors, honeypot weaknesses in the classification manifold that attract attackers searching for adversarial examples. Attackers' optimization algorithms gravitate towards trapdoors, leading them to produce attacks similar to trapdoors in the feature space. Our defense then identifies attacks by comparing neuron activation signatures of inputs to those of trapdoors. In this paper, we introduce trapdoors and describe an implementation of a trapdoor-enabled defense. First, we analytically prove that trapdoors shape the computation of adversarial attacks so that attack inputs will have feature representations very similar to those of trapdoors. Second, we experimentally show that trapdoor-protected models can detect, with high accuracy, adversarial examples generated by state-of-the-art attacks (PGD, optimization-based CW, Elastic Net, BPDA), with negligible impact on normal classification. These results generalize across classification domains, including image, facial, and traffic-sign recognition. We also present significant results measuring trapdoors' robustness against customized adaptive attacks (countermeasures).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao 等CCS 2021 · 被引用 108 次
- Handcrafted Backdoors in Deep Neural NetworksSanghyun Hong, Nicholas Carlini, Alexey KurakinNeurIPS 2022 · 被引用 105 次
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsShawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu 等S&P 2024 · 被引用 102 次
- Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient DescentOliver Bryniarski, Nabeel Hingun, Pedro Pachuca, Vincent Wang 等ICLR 2022 · 被引用 43 次
- Attack as defense: characterizing adversarial examples using robustnessZhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang 等ISSTA 2021 · 被引用 34 次
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
相关 Paper
- NeuroPots: Realtime Proactive Defense against Bit-Flip Attacks in Neural NetworksQi Liu, Jieming Yin, Wujie Wen, Chengmo Yang 等USENIX Security 2023
- Feature-Indistinguishable Attack to Circumvent Trapdoor-Enabled DefenseChaoxiang He, Bin Benjamin Zhu, Xiaojing Ma, Hai Jin 等CCS 2021 · 被引用 4 次
- Trap-MID: Trapdoor-based Defense against Model Inversion AttacksZhenTing Liu, ShangTse ChenNeurIPS 2024 · 被引用 12 次
- Backdoor Defense via Enhanced Splitting and Trap IsolationHongrui Yu, Lu Qi, Wanyu Lin, Jian Chen 等ICCV 2025 · 被引用 5 次
- TrojanFlow: A Neural Backdoor Attack to Deep Learning-based Network Traffic ClassifiersRui Ning, Chunsheng Xin, Hongyi WuINFOCOM 2022 · 被引用 27 次
