AdvMind: Inferring Adversary Intent of Black-Box Attacks
Ren Pang, Xinyang Zhang, Shouling Ji, Xiapu Luo, Ting Wang
Abstract
Deep neural networks (DNNs) are inherently susceptible to adversarial attacks even under black-box settings, in which the adversary only has query access to the target models. In practice, while it may be possible to effectively detect such attacks (e.g., observing massive similar but non-identical queries), it is often challenging to exactly infer the adversary intent (e.g., the target class of the adversarial example the adversary attempts to craft) especially during early stages of the attacks, which is crucial for performing effective deterrence and remediation of the threats in many scenarios.
In this paper, we present AdvMind, a new class of estimation models that infer the adversary intent of black-box adversarial attacks in a robust and prompt manner. Specifically, to achieve robust detection, AdvMind accounts for the adversary adaptiveness such that her attempt to conceal the target will significantly increase the attack cost (e.g., in terms of the number of queries); to achieve prompt detection, AdvMind proactively synthesizes plausible query results to solicit subsequent queries from the adversary that maximally expose her intent. Through extensive empirical evaluation on benchmark datasets and state-of-the-art black-box attacks, we demonstrate that on average AdvMind detects the adversary intent with over 75% accuracy after observing less than 3 query batches and meanwhile increases the cost of adaptive attacks by over 60%. We further discuss the possible synergy between AdvMind and other defense methods against black-box adversarial attacks, pointing to several promising research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Too Good to Be Safe: Tricking Lane Detection in Autonomous Driving with Crafted PerturbationsPengfei Jing, Qiyi Tang, Yuefeng Du, Lei Xue et al.USENIX Security 2021 · 79 citations
- Random Noise Defense Against Query-Based Black-Box AttacksZeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan WuNeurIPS 2021 · 78 citations
- Robust Spatiotemporal Traffic Forecasting with Reinforced Dynamic Adversarial TrainingFan Liu, Weijia Zhang, Hao LiuKDD 2023 · 15 citations
- Stateful Defenses for Machine Learning Models Are Not Yet Secure Against Black-box AttacksRyan Feng, Ashish Hooda, Neal Mangaokar, Kassem Fawaz et al.CCS 2023 · 10 citations
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong et al.ICLR 2024 · 3 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- DEEPSEC: A Uniform Platform for Security Analysis of Deep Learning ModelXiang Ling, Shouling Ji, Jiaxu Zou, Jiannan Wang et al.S&P 2019 · 147 citations
Related papers
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- AdvHunter: Detecting Adversarial Perturbations in Black-Box Neural Networks through Hardware Performance CountersManaar Alam, Michail ManiatakosDAC 2024 · 3 citations
- DeepRover: A Query-Efficient Blackbox Attack for Deep Neural NetworksFuyuan Zhang, Xinwen Hu, Lei Ma, Jianjun ZhaoFSE 2023 · 7 citations
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang et al.USENIX Security 2022
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang et al.ICCV 2021 · 128 citations
