Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial Attacks
Nguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong, Khoa D. Doan
摘要
Recent works have shown that deep neural networks are vulnerable to adversarial examples that find samples close to the original image but can make the model misclassify. Even with access only to the model's output, an attacker can employ black-box attacks to generate such adversarial examples. In this work, we propose a simple and lightweight defense against black-box attacks by adding random noise to hidden features at intermediate layers of the model at inference time. Our theoretical analysis confirms that this method effectively enhances the model's resilience against both score-based and decision-based black-box attacks. Importantly, our defense does not necessitate adversarial training and has minimal impact on accuracy, rendering it applicable to any pre-trained model. Our analysis also reveals the significance of selectively adding noise to different parts of the model based on the gradient of the adversarial objective function, which can be varied during the attack. We demonstrate the robustness of our defense against multiple black-box attacks through extensive empirical experiments involving diverse models with various architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- RayS: A Ray Searching Method for Hard-label Adversarial AttackJinghui Chen, Quanquan GuKDD 2020 · 被引用 108 次
- Sign Bits Are All You Need for Black-Box AttacksAbdullah Al-Dujaili, Una-May O'ReillyICLR 2020 · 被引用 93 次
- Random Noise Defense Against Query-Based Black-Box AttacksZeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan WuNeurIPS 2021 · 被引用 78 次
相关 Paper
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang 等USENIX Security 2022
- LEA2: A Lightweight Ensemble Adversarial Attack via Non-overlapping Vulnerable Frequency RegionsYaguan Qian, Shuke He, Chenyu Zhao, Jiaqiang Sha 等ICCV 2023 · 被引用 26 次
- Towards Lightweight Black-Box Attack Against Deep Neural NetworksChenghao Sun, Yonggang Zhang, Chaoqun Wan, Qizhou Wang 等NeurIPS 2022 · 被引用 12 次
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 被引用 282 次
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke 等ICCV 2019 · 被引用 160 次
