Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis
Jeonghwan Park, Niall McLaughlin, Ihsen Alouani
Abstract
Adversarial attacks remain a significant threat that can jeopardize the integrity of Machine Learning (ML) models. In particular, query-based black-box attacks can generate malicious noise without having access to the victim model's architecture, making them practical in real-world contexts. The community has proposed several defenses against adversarial attacks, only to be broken by more advanced and adaptive attack strategies. In this paper, we propose a framework that detects if an adversarial noise instance is being generated. Unlike existing stateful defenses that detect adversarial noise generation by monitoring the input space, our approach learns adversarial patterns in the input update similarity space. In fact, we propose to observe a new metric called Delta Similarity (DS), which we show it captures more efficiently the adversarial behavior. We evaluate our approach against 8 state-of-the-art attacks, including adaptive attacks, where the adversary is aware of the defense and tries to evade detection. We find that our approach is significantly more robust than existing defenses both in terms of specificity and sensitivity. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning CompilersSimin Chen, Jinjun Peng, Yixin He, Junfeng Yang et al.S&P 2026 · 11 citations
- RemedyGS: Defend 3D Gaussian Splatting Against Computation Cost AttacksYanping Li, Zhening Liu, Zijian Li, Zehong Lin et al.CVPR 2026 · 7 citations
Builds on8
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Sign-OPT: A Query-Efficient Hard-label Adversarial AttackMinhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen et al.ICLR 2020 · 256 citations
- Defensive approximation: securing CNNs using approximate computingAmira Guesmi, Ihsen Alouani, Khaled N. Khasawneh, Mouna Baklouti et al.ASPLOS 2021 · 46 citations
- Stateful Defenses for Machine Learning Models Are Not Yet Secure Against Black-box AttacksRyan Feng, Ashish Hooda, Neal Mangaokar, Kassem Fawaz et al.CCS 2023 · 10 citations
- QEBA: Query-Efficient Boundary-Based Blackbox AttackHuichen Li, Xiaojun Xu, Xiaolu Zhang, Shuang Yang et al.CVPR 2020
Related papers
- Query Provenance Analysis: Efficient and Robust Defense Against Query-Based Black-Box AttacksShaofei Li, Ziqi Zhang, Haomin Jia, Yao Guo et al.S&P 2025
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang et al.USENIX Security 2022
- Attack To Defend: Exploiting Adversarial Attacks for Detecting Poisoned ModelsSamar Fares, Karthik NandakumarCVPR 2024
- Learning to Generate Noise for Multi-Attack RobustnessDivyam Madaan, Jinwoo Shin, Sung Ju HwangICML 2021 · 31 citations
- Attack as defense: characterizing adversarial examples using robustnessZhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang et al.ISSTA 2021 · 34 citations
