Feature-Indistinguishable Attack to Circumvent Trapdoor-Enabled Defense
Chaoxiang He, Bin Benjamin Zhu, Xiaojing Ma, Hai Jin, Shengshan Hu
摘要
Deep neural networks (DNNs) are vulnerable to adversarial attacks. A great effort has been directed to developing effective defenses against adversarial attacks and finding vulnerabilities of proposed defenses. A recently proposed defense called Trapdoor-enabled Detection (TeD) deliberately injects trapdoors into DNN models to trap and detect adversarial examples targeting categories protected by TeD. TeD can effectively detect existing state-of-the-art adversarial attacks. In this paper, we propose a novel black-box adversarial attack on TeD, called Feature-Indistinguishable Attack (FIA). It circumvents TeD by crafting adversarial examples indistinguishable in the feature (i.e., neuron-activation) space from benign examples in the target category. To achieve this goal, FIA jointly minimizes the distance to the expectation of feature representations of benign samples in the target category and maximizes the distances to positive adversarial examples generated to query TeD in the preparation phase. A constraint is used to ensure that the feature vector of a generated adversarial example is within the distribution of feature vectors of benign examples in the target category. Our extensive empirical evaluation with different configurations and variants of TeD indicates that our proposed FIA can effectively circumvent TeD. FIA opens a door for developing much more powerful adversarial attacks. The FIA code is available at: https://github.com/CGCL-codes/FeatureIndistinguishableAttack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsShawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng 等CCS 2022 · 被引用 9 次
- What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic ContradictionYijun Yang, Ruiyuan Gao, Yu Li, Qiuxia Lai 等NDSS 2022
- DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable DefensesChaoxiang He, Xiaojing Ma, Bin B. Zhu, Yimiao Zeng 等NDSS 2024
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
相关 Paper
- Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural NetworksShawn Shan, Emily Wenger, Bolun Wang, Bo Li 等CCS 2020 · 被引用 67 次
- NeuroPots: Realtime Proactive Defense against Bit-Flip Attacks in Neural NetworksQi Liu, Jieming Yin, Wujie Wen, Chengmo Yang 等USENIX Security 2023
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 被引用 191 次
- Generating Distributional Adversarial Examples to Evade Statistical DetectorsYigitcan Kaya, Muhammad Bilal Zafar, Sergül Aydöre, Nathalie Rauschmayr 等ICML 2022 · 被引用 7 次
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang 等ICCV 2021 · 被引用 128 次
