Stateful Defenses for Machine Learning Models Are Not Yet Secure Against Black-box Attacks
Ryan Feng, Ashish Hooda, Neal Mangaokar, Kassem Fawaz, Somesh Jha, Atul Prakash
摘要
Recent work has proposed stateful defense models (SDMs) as a compelling strategy to defend against a black-box attacker who only has query access to the model, as is common for online machine learning platforms. Such stateful defenses aim to defend against black-box attacks by tracking the query history and detecting and rejecting queries that are "similar" and thus preventing black-box attacks from finding useful gradients and making progress towards finding adversarial attacks within a reasonable query budget. Recent SDMs (e.g., Blacklight and PIHA) have shown remarkable success in defending against state-of-the-art black-box attacks. In this paper, we show that SDMs are highly vulnerable to a new class of adaptive black-box attacks. We propose a novel adaptive black-box attack strategy called Oracle-guided Adaptive Rejection Sampling (OARS) that involves two stages: (1) use initial query patterns to infer key properties about an SDM's defense; and, (2) leverage those extracted properties to design subsequent query patterns to evade the SDM's defense while making progress towards finding adversarial inputs. OARS is broadly applicable as an enhancement to existing black-box attacks -we show how to apply the strategy to enhance six common black-box attacks to be more effective against current class of SDMs. For example, OARS-enhanced versions of black-box attacks improved attack success rate against recent stateful defenses from almost 0% to to almost 100% for multiple datasets within reasonable query budgets. CCS CONCEPTS • Computing methodologies → Machine learning; • Security and privacy; † Denotes equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- SoK: The Pitfalls of Deep Reinforcement Learning for CybersecurityShae McFadden, Myles Foley, Elizabeth Bates, Ilias Tsingenopoulos 等USENIX Security 2026 · 被引用 7 次
- AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt TuningXin Wang, Kai Chen, Xingjun Ma, Zhineng Chen 等ACM MM 2024 · 被引用 6 次
- PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMsYuchen Yang, Yiming Li, Hongwei Yao, Enhao Huang 等S&P 2026 · 被引用 5 次
- Query Provenance Analysis: Efficient and Robust Defense Against Query-Based Black-Box AttacksShaofei Li, Ziqi Zhang, Haomin Jia, Yao Guo 等S&P 2025
它引用的顶会 Paper10
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 被引用 797 次
- Deepfake Videos in the Wild: Analysis and DetectionJiameng Pu, Neal Mangaokar, Lauren Kelly, Parantapa Bhattacharya 等WWW 2021 · 被引用 59 次
- Adversarial Attack on Attackers: Post-Process to Mitigate Black-Box Score-Based Query AttacksSizhe Chen, Zhehao Huang, Qinghua Tao, Yingwen Wu 等NeurIPS 2022 · 被引用 36 次
相关 Paper
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang 等USENIX Security 2022
- Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update AnalysisJeonghwan Park, Niall McLaughlin, Ihsen AlouaniCVPR 2025
- Embracing Adaptation: An Effective Dynamic Defense Strategy Against Adversarial ExamplesShenglin Yin, Kelu Yao, Zhen Xiao, Jieyi LongACM MM 2024
- Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable ConfidenceHanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba 等CCS 2024 · 被引用 3 次
- Low-Cost Hard-Label Adversarial Attack with Theoretical FoundationsJun Liu, Leo Yu Zhang, Fengpeng Li, Isao Echizen 等USENIX Security 2026
