Consistency-Sensitivity Guided Ensemble Black-Box Adversarial Attacks in Low-Dimensional Spaces
Jianhe Yuan, Zhihai He
Abstract
Black-box attacks aim to generate adversarial noise to fail the victim deep neural network in the black box. The central task in black-box attack method design is to estimate and characterize the victim model in the high-dimensional model space based on feedback results of queries submitted to the victim network. The central performance goal is to minimize the number of queries needed for successful at-tack. Existing attack methods directly search and refine the adversarial noise in an extremely high-dimensional space, requiring hundreds or even thousands queries to the victim network. To address this challenge, we propose to explore a consistency and sensitivity guided ensemble attack (CSEA) method in a low-dimensional space. Specifically, we estimate the victim model in the black box using a learned linear composition of an ensemble of surrogate models with diversified network structures. Using random block masks on the input image, these surrogate models jointly construct and submit randomized and sparsified queries to the victim model. Based on these query results and guided by a consistency constraint, the surrogate models can be trained using a very small number of queries such that their learned composition is able to accurately approximate the victim model in the high-dimensional space. The randomized and sparsified queries also provide important information for us to construct an attack sensitivity map for the input image, with which the adversarial attack can be locally refined to further increase its success rate. Our extensive experimental results demonstrate that our proposed approach significantly reduces the number of queries to the victim network while maintaining very high success rates, outperforming existing black-box attack methods by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Blackbox Attacks via Surrogate Ensemble SearchZikui Cai, Chengyu Song, Srikanth V. Krishnamurthy, Amit Roy-Chowdhury et al.NeurIPS 2022 · 32 citations
- Learning Black-Box Attackers with Transferable Priors and Query FeedbackJiancheng Yang, Yangzhou Jiang, Xiaoyang Huang, Bingbing Ni et al.NeurIPS 2020 · 96 citations
- BayesOpt Adversarial AttackBinxin Ru, Adam D. Cobb, Arno Blaas, Yarin GalICLR 2020 · 85 citations
- Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget RegimesSatya Narayan Shukla, Anit Kumar Sahu, Devin Willmott, J. Zico KolterKDD 2021 · 24 citations
- Enhancing Adversarial Transferability with Checkpoints of a Single Model's TrainingShixin Li, Chaoxiang He, Xiaojing Ma, Bin Benjamin Zhu et al.CVPR 2025
