Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box Attacks
Huiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang, Haitao Zheng, Ben Y. Zhao
摘要
Deep learning systems are known to be vulnerable to adversarial examples. In particular, query-based black-box attacks do not require knowledge of the deep learning model, but can compute adversarial examples over the network by submitting queries and inspecting returns. Recent work largely improves the efficiency of those attacks, demonstrating their practicality on today's ML-as-a-service platforms. We propose Blacklight, a new defense against query-based black-box adversarial attacks. The fundamental insight driving our design is that, to compute adversarial examples, these attacks perform iterative optimization over the network, producing image queries highly similar in the input space. Blacklight detects query-based black-box attacks by detecting highly similar queries, using an efficient similarity engine operating on probabilistic content fingerprints. We evaluate Blacklight against eight state-of-the-art attacks, across a variety of models and image classification tasks. Blacklight identifies them all, often after only a handful of queries. By rejecting all detected queries, Blacklight prevents any attack to complete, even when attackers persist to submit queries after account ban or query rejection. Blacklight is also robust against several powerful countermeasures, including an optimal black-box attack that approximates white-box attacks in efficiency. Finally, we illustrate how Blacklight generalizes to other domains like text classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsDaizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou 等NeurIPS 2024 · 被引用 51 次
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan 等NeurIPS 2023 · 被引用 26 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- OSLO: One-Shot Label-Only Membership Inference AttacksYuefeng Peng, Jaechul Roh, Subhransu Maji, Amir HoumansadrNeurIPS 2024 · 被引用 17 次
- Efficient Query-Based Attack against ML-Based Android Malware Detection under Zero Knowledge SettingPing He, Yifan Xia, Xuhong Zhang, Shouling JiCCS 2023 · 被引用 13 次
它引用的顶会 Paper19
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
相关 Paper
- Query Provenance Analysis: Efficient and Robust Defense Against Query-Based Black-Box AttacksShaofei Li, Ziqi Zhang, Haomin Jia, Yao Guo 等S&P 2025
- Black-Box Adversarial Attack with Transferable Model-based EmbeddingZhichao Huang, Tong ZhangICLR 2020 · 被引用 131 次
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong 等ICLR 2024 · 被引用 3 次
- Stateful Defenses for Machine Learning Models Are Not Yet Secure Against Black-box AttacksRyan Feng, Ashish Hooda, Neal Mangaokar, Kassem Fawaz 等CCS 2023 · 被引用 10 次
- AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt TuningXin Wang, Kai Chen, Xingjun Ma, Zhineng Chen 等ACM MM 2024 · 被引用 6 次
