What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
Yijun Yang, Ruiyuan Gao, Yu Li, Qiuxia Lai, Qiang Xu
Abstract
Adversarial examples (AEs) pose severe threats to the applications of deep neural networks (DNNs) to safety-critical domains, e.g., autonomous driving. While there has been a vast body of AE defense solutions, to the best of our knowledge, they all suffer from some weaknesses, e.g., defending against only a subset of AEs or causing a relatively high accuracy loss for legitimate inputs. Moreover, most existing solutions cannot defend against adaptive attacks, wherein attackers are knowledgeable about the defense mechanisms and craft AEs accordingly. In this paper, we propose a novel AE detection framework based on the very nature of AEs, i.e., their semantic information is inconsistent with the discriminative features extracted by the target DNN model. To be specific, the proposed solution, namely ContraNet, models such contradiction by first taking both the input and the inference result to a generator to obtain a synthetic output and then comparing it against the original input. For legitimate inputs that are correctly inferred, the synthetic output tries to reconstruct the input. On the contrary, for AEs, instead of reconstructing the input, the synthetic output would be created to conform to the wrong label whenever possible. Consequently, by measuring the distance between the input and the synthetic output with metric learning, we can differentiate AEs from legitimate inputs. We perform comprehensive evaluations under various AE attack scenarios, and experimental results show that ContraNet outperforms existing solutions by a large margin, especially under adaptive attacks. Moreover, our analysis shows that successful AEs that can bypass ContraNet tend to have much-weakened adversarial semantics. We have also shown that ContraNet can be easily combined with adversarial training techniques to achieve further improved AE defense capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c76ccf7-1bd7-43c7-89db-80367633ece2Cited by top-tier papers5
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 22 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningZhiyuan He, Yijun Yang, Pin-Yu Chen, Qiang Xu et al.ICML 2024 · 11 citations
- Towards Targeted Obfuscation of Adversarial Unsafe Images using Reconstruction and Counterfactual Super Region Attribution ExplainabilityMazal Bethany, Andrew Seong, Samuel Henrique Silva, Nicole Beebe et al.USENIX Security 2023
- Constructive Noise Defeats Adversarial Noise: Adversarial Example Detection for Commercial DNN ServicesMeng Shen, Jiangyuan Bi, Hao Yu, Zhenming Bai et al.NDSS 2026
Builds on22
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
Related papers
- Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform DomainJinyu Tian, Jiantao Zhou, Yuanman Li, Jia DuanAAAI 2021 · 72 citations
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative ModelsTong Che, Xiaofeng Liu, Site Li, Yubin Ge et al.AAAI 2021 · 54 citations
- Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyXiyue Zhang, Xiaofei Xie, Lei Ma, Xiaoning Du et al.ICSE 2020 · 69 citations
- AdvIT: Adversarial Frames Identifier Based on Temporal Consistency in VideosChaowei Xiao, Ruizhi Deng, Bo Li, Taesung Lee et al.ICCV 2019 · 64 citations
- Fooling the Eyes of Autonomous Vehicles: Robust Physical Adversarial Examples Against Traffic Sign Recognition SystemsWei Jia, Zhaojun Lu, Haichun Zhang, Zhenglin Liu et al.NDSS 2022
