Multi-granularity Textual Adversarial Attack with Behavior Cloning
Yangyi Chen, Jin Su, Wei Wei
Abstract
Recently, the textual adversarial attack models become increasingly popular due to their successful in estimating the robustness of NLP models. However, existing works have obvious deficiencies. (1) They usually consider only a single granularity of modification strategies (e.g. word-level or sentence-level), which is insufficient to explore the holistic textual space for generation; (2) They need to query victim models hundreds of times to make a successful attack, which is highly inefficient in practice. To address such problems, in this paper we propose MAYA, a Multi-grAnularitY Attack model to effectively generate high-quality adversarial samples with fewer queries to victim models. Furthermore, we propose a reinforcement-learning based method to train a multi-granularity attack agent through behavior cloning with the expert knowledge from our MAYA algorithm to further reduce the query times. Additionally, we also adapt the agent to attack blackbox models that only output labels without confidence scores. We conduct comprehensive experiments to evaluate our attack models by attacking BiLSTM, BERT and RoBERTa in two different black-box attack settings and three benchmark datasets. Experimental results show that our models achieve overall better attacking performance and produce more fluent and grammatical adversarial samples compared to baseline models. Besides, our adversarial attack agent significantly reduces the query times in both attack settings. Our codes are released at https://github. com/Yangyi-Chen/MAYA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79f203b-5639-4e17-91aa-72b9a2ce8445Cited by top-tier papers6
- Multi-granular Adversarial Attacks against Black-box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2024 · 17 citations
- White-Box Multi-Objective Adversarial Attack on Dialogue GenerationYufei Li, Zexin Li, Yingfan Gao, Cong LiuACL 2023 · 12 citations
- TASA: Deceiving Question Answering Models by Twin Answer Sentences AttackYu Cao, Dianqi Li, Meng Fang, Tianyi Zhou et al.EMNLP 2022 · 10 citations
- Are LLM-Enhanced Graph Neural Networks Robust Against Poisoning Attacks?Yuhang Ma, Jie Wang, Zheng YanS&P 2026 · 4 citations
- Generative Adversarial Training with Perturbed Token Detection for Model RobustnessJiahao Zhao, Wenji MaoEMNLP 2023 · 3 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 128 citations
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 92 citations
Related papers
- Experience Speaks Louder: Black-box Hard-label Adversarial Attack through Reinforcement LearningYilun Jin, Kun Zhu, Feng Tang, Jiaxuan Shi et al.KDD 2025 · 1 citation
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted AttackBoxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen et al.EMNLP 2020 · 55 citations
- Revisiting Character-level Adversarial Attacks for Language ModelsElías Abad-Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos et al.ICML 2024 · 1 citation
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on TextHan Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang et al.NeurIPS 2023 · 32 citations
- Multi-task Adversarial Attacks against Black-box Model with Few-shot QueriesWenqiang Wang, Yan Xiao, Hao Lin, Yangshijie Zhang et al.ACL 2025 · 1 citation
