Sign-OPT: A Query-Efficient Hard-label Adversarial Attack
Minhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen, Sijia Liu, Cho-Jui Hsieh
Abstract
We study the most practical problem setup for evaluating adversarial robustness of a machine learning system with limited access: the hard-label black-box attack setting for generating adversarial examples, where limited model queries are allowed and only the decision is provided to a queried data input. Several algorithms have been proposed for this problem but they typically require huge amount (>20,000) of queries for attacking one example. Among them, one of the state-of-the-art approaches (Cheng et al., 2019) showed that hard-label attack can be modeled as an optimization problem where the objective function can be evaluated by binary search with additional model queries, thereby a zeroth order optimization algorithm can be applied. In this paper, we adopt the same optimization formulation but propose to directly estimate the sign of gradient at any direction instead of the gradient itself, which enjoys the benefit of single query. Using this single query oracle for retrieving sign of directional derivative, we develop a novel query-efficient Sign-OPT approach for hard-label black-box attack. We provide a convergence analysis of the new algorithm and conduct experiments on several models on MNIST, CIFAR-10 and ImageNet. We find that Sign-OPT attack consistently requires 5X to 10X fewer queries when compared to the current state-of-the-art approaches, and usually converges to an adversarial example with smaller perturbation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5ef5084-c765-40d6-9d15-a112316bbcf0Cited by top-tier papers59
- Boosting the Transferability of Adversarial Attacks with Reverse Adversarial PerturbationZeyu Qin, Yanbo Fan, Yi Liu, Li Shen et al.NeurIPS 2022 · 135 citations
- Transfer Learning without Knowing: Reprogramming Black-box Machine Learning Models with Scarce Data and Limited ResourcesYun-Yun Tsai, Pin-Yu Chen, Tsung-Yi HoICML 2020 · 115 citations
- Diversity can be Transferred: Output Diversification for White- and Black-box AttacksYusuke Tashiro, Yang Song, Stefano ErmonNeurIPS 2020 · 114 citations
- RayS: A Ray Searching Method for Hard-label Adversarial AttackJinghui Chen, Quanquan GuKDD 2020 · 108 citations
- AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value AnalysisJunfeng Guo, Ang Li, Cong LiuICLR 2022 · 92 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Guessing Smart: Biased Sampling for Efficient Black-Box Adversarial AttacksThomas Brunner, Frederik Diehl, Michael Truong-Le, Alois C. KnollICCV 2019 · 127 citations
Related papers
- Sign Bits Are All You Need for Black-Box AttacksAbdullah Al-Dujaili, Una-May O'ReillyICLR 2020 · 93 citations
- Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based PriorsChen Ma, Xinjie Xu, Shuyu Cheng, Qi XuanICLR 2025
- Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget RegimesSatya Narayan Shukla, Anit Kumar Sahu, Devin Willmott, J. Zico KolterKDD 2021 · 24 citations
- Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label AttacksXinjie Xu, Shuyu Cheng, Dongwei Xu, Qi Xuan et al.AAAI 2026
- Policy-Driven Attack: Learning to Query for Hard-label Black-box Adversarial ExamplesZiang Yan, Yiwen Guo, Jian Liang, Changshui ZhangICLR 2021 · 19 citations
