Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial Attacks
Nguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong, Khoa D. Doan
Abstract
Recent works have shown that deep neural networks are vulnerable to adversarial examples that find samples close to the original image but can make the model misclassify. Even with access only to the model's output, an attacker can employ black-box attacks to generate such adversarial examples. In this work, we propose a simple and lightweight defense against black-box attacks by adding random noise to hidden features at intermediate layers of the model at inference time. Our theoretical analysis confirms that this method effectively enhances the model's resilience against both score-based and decision-based black-box attacks. Importantly, our defense does not necessitate adversarial training and has minimal impact on accuracy, rendering it applicable to any pre-trained model. Our analysis also reveals the significance of selectively adding noise to different parts of the model based on the gradient of the adversarial objective function, which can be varied during the attack. We demonstrate the robustness of our defense against multiple black-box attacks through extensive empirical experiments involving diverse models with various architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b2dc4b0-8e2d-4260-8d52-a7c486658f32Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- RayS: A Ray Searching Method for Hard-label Adversarial AttackJinghui Chen, Quanquan GuKDD 2020 · 108 citations
- Sign Bits Are All You Need for Black-Box AttacksAbdullah Al-Dujaili, Una-May O'ReillyICLR 2020 · 93 citations
- Random Noise Defense Against Query-Based Black-Box AttacksZeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan WuNeurIPS 2021 · 78 citations
Related papers
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang et al.USENIX Security 2022
- LEA2: A Lightweight Ensemble Adversarial Attack via Non-overlapping Vulnerable Frequency RegionsYaguan Qian, Shuke He, Chenyu Zhao, Jiaqiang Sha et al.ICCV 2023 · 26 citations
- Towards Lightweight Black-Box Attack Against Deep Neural NetworksChenghao Sun, Yonggang Zhang, Chaoqun Wan, Qizhou Wang et al.NeurIPS 2022 · 12 citations
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 282 citations
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke et al.ICCV 2019 · 160 citations
