Attack as defense: characterizing adversarial examples using robustness
Zhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang, Fu Song, Jun Sun
摘要
As a new programming paradigm, deep learning has expanded its application to many real-world problems. At the same time, deep learning based software are found to be vulnerable to adversarial attacks. Though various defense mechanisms have been proposed to improve robustness of deep learning software, many of them are ineffective against adaptive attacks. In this work, we propose a novel characterization to distinguish adversarial examples from benign ones based on the observation that adversarial examples are significantly less robust than benign ones. As existing robustness measurement does not scale to large networks, we propose a novel defense framework, named attack as defense (A 2 D), to detect adversarial examples by effectively evaluating an example's robustness. A 2 D uses the cost of attacking an input for robustness evaluation and identifies those less robust examples as adversarial since less robust examples are easier to attack. Extensive experiment results on MNIST, CIFAR10 and ImageNet show that A 2 D is more effective than recent promising approaches. We also evaluate our defence against potential adaptive attacks and show that A 2 D is effective in defending carefully designed adaptive attacks, e.g., the attack success rate drops to 0% on CIFAR10.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 被引用 91 次
- QVIP: An ILP-based Formal Verification Approach for Quantized Neural NetworksYedi Zhang, Zhe Zhao, Guangke Chen, Fu Song 等ASE 2022 · 被引用 21 次
- QEBVerif: Quantization Error Bound Verification of Neural NetworksYedi Zhang, Fu Song, Jun SunCAV 2023 · 被引用 17 次
- DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsLongtian Wang, Xiaofei Xie, Xiaoning Du, Meng Tian 等FSE 2023 · 被引用 15 次
- Toward Improving the Robustness of Deep Learning Models via Model TransformationYingyi Zhang, Zan Wang, Jiajun Jiang, Hanmo You 等ASE 2022 · 被引用 7 次
它引用的顶会 Paper15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
相关 Paper
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja 等NeurIPS 2021 · 被引用 22 次
- A Unified Framework for Detecting Audio Adversarial ExamplesXia Du, Chi-Man Pun, Zheng ZhangACM MM 2020 · 被引用 18 次
- SuperDeepFool: a new fast and accurate minimal adversarial attackAlireza Abdollahpour, Mahed Abroshan, Seyed-Mohsen Moosavi-DezfooliNeurIPS 2024 · 被引用 11 次
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee 等NDSS 2019 · 被引用 283 次
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 被引用 23 次
