Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
Hanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba, Yuan Hong
Abstract
Black-box adversarial attacks have demonstrated strong potential to compromise machine learning models by iteratively querying the target model or leveraging transferability from a local surrogate model. Recently, such attacks can be effectively mitigated by state-of-the-art (SOTA) defenses, e.g., detection via the pattern of sequential queries, or injecting noise into the model. To our best knowledge, we take the first step to study a new paradigm of black-box attacks with provable guarantees -certifiable black-box attacks that can guarantee the attack success probability (ASP) of adversarial examples before querying over the target model. This new black-box attack unveils significant vulnerabilities of machine learning models, compared to traditional empirical black-box attacks, e.g., breaking strong SOTA defenses with provable confidence, constructing a space of (infinite) adversarial examples with high ASP, and the ASP of the generated adversarial examples is theoretically guaranteed without verification/queries over the target model. Specifically, we establish a novel theoretical foundation for ensuring the ASP of the black-box attack with randomized adversarial examples (AEs). Then, we propose several novel techniques to craft the randomized AEs while reducing the perturbation size for better imperceptibility. Finally, we have comprehensively evaluated the certifiable black-box attacks on the CIFAR10/100, ImageNet, and LibriSpeech datasets, while benchmarking with 16 SOTA black-box attacks, against various SOTA defenses in the domains of computer vision and speech recognition. Both theoretical and experimental results have validated the significance of the proposed attack. 1 CCS CONCEPTS • Security and privacy → Formal security models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9480bd64-b865-4a0f-989d-497bfb9e04f6Cited by top-tier papers5
- Learning Robust and Privacy-Preserving Representations via Information TheoryBinghui Zhang, Sayedeh Leila Noorbakhsh, Yun Dong, Yuan Hong et al.AAAI 2025 · 4 citations
- Adversarial Attack on Black-Box Multi-Agent by Adaptive PerturbationJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie et al.AAAI 2026 · 1 citation
- PLRV-O: Advancing Differentially Private Deep Learning via Privacy Loss Random Variable OptimizationQin Yang, Nicholas Stout, Meisam Mohammady, Han Wang et al.CCS 2025
- What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text AttacksQin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang et al.USENIX Security 2026
- Low-Cost Hard-Label Adversarial Attack with Theoretical FoundationsJun Liu, Leo Yu Zhang, Fengpeng Li, Isao Echizen et al.USENIX Security 2026
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
Related papers
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 31 citations
- Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition SystemsZheng Fang, Tao Wang, Lingchen Zhao, Shenyi Zhang et al.CCS 2024 · 11 citations
- Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box AttacksHuiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang et al.USENIX Security 2022
- Towards Multiple Black-boxes Attack via Adversarial Example Generation NetworkMingxing Duan, Kenli Li, Lingxi Xie, Qi Tian et al.ACM MM 2021 · 21 citations
- Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function PriorShuyu Cheng, Yibo Miao, Yinpeng Dong, Xiao Yang et al.ICML 2024 · 15 citations
