LeapAttack: Hard-Label Adversarial Attack on Text via Gradient-Based Optimization
Muchao Ye, Jinghui Chen, Chenglin Miao, Ting Wang, Fenglong Ma
Abstract
Generating text adversarial examples in the hard-label setting is a more realistic and challenging black-box adversarial attack problem, whose challenge comes from the fact that gradient cannot be directly calculated from discrete word replacements. Consequently, the effectiveness of gradient-based methods for this problem still awaits improvement. In this paper, we propose a gradient-based optimization method named LeapAttack to craft high-quality text adversarial examples in the hard-label setting. To specify, LeapAttack employs the word embedding space to characterize the semantic deviation between the two words of each perturbed substitution by their difference vector. Facilitated by this expression, LeapAttack gradually updates the perturbation direction and constructs adversarial examples in an iterative round trip: firstly, the gradient is estimated by transforming randomly sampled word candidates to continuous difference vectors after moving the current adversarial example near the decision boundary; secondly, the estimated gradient is mapped back to a new substitution word based on the cosine similarity metric. Extensive experimental results show that in the general case LeapAttack can efficiently generate high-quality text adversarial examples with the highest semantic similarity and the lowest perturbation rate in the hard-label setting. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10cdb489-a799-49b8-a05e-88f25e7449beCited by top-tier papers4
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on TextHan Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang et al.NeurIPS 2023 · 32 citations
- LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial AttackHai Zhu, Qingyang Zhao, Weiwei Shang, Yuren Wu et al.AAAI 2024 · 19 citations
- UniT: A Unified Look at Certified Robust Training against Text Adversarial PerturbationMuchao Ye, Ziyi Yin, Tianrong Zhang, Tianyu Du et al.NeurIPS 2023 · 3 citations
- RAt: Injecting Implicit Bias for Text-To-Image Prompt Refinement ModelsZiyi Kou, Shichao Pei, Meng Jiang, Xiangliang ZhangEMNLP 2024 · 1 citation
Builds on7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Sign-OPT: A Query-Efficient Hard-label Adversarial AttackMinhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen et al.ICLR 2020 · 256 citations
Related papers
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu et al.AAAI 2023 · 31 citations
- TextHoaxer: Budgeted Hard-Label Adversarial Attacks on TextMuchao Ye, Chenglin Miao, Ting Wang, Fenglong MaAAAI 2022 · 57 citations
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 128 citations
- PAT: Geometry-Aware Hard-Label Black-Box Adversarial Attacks on TextMuchao Ye, Jinghui Chen, Chenglin Miao, Han Liu et al.KDD 2023 · 8 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
