Word-level Textual Adversarial Attacking as Combinatorial Optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, Maosong Sun
摘要
Adversarial attacks are carried out to reveal the vulnerability of deep neural networks. Textual adversarial attacking is challenging because text is discrete and a small perturbation can bring significant change to the original input. Word-level attacking, which can be regarded as a combinatorial optimization problem, is a well-studied class of textual attack methods. However, existing word-level attack models are far from perfect, largely because unsuitable search space reduction methods and inefficient optimization algorithms are employed. In this paper, we propose a novel attack model, which incorporates the sememebased word substitution method and particle swarm optimization-based search algorithm to solve the two problems separately. We conduct exhaustive experiments to evaluate our attack model by attacking BiLSTM and BERT on three benchmark datasets. Experimental results demonstrate that our model consistently achieves much higher attack success rates and crafts more high-quality adversarial examples as compared to baseline methods. Also, further experiments show our model has higher transferability and can bring more robustness enhancement to victim models by adversarial training. All the code and data of this paper can be obtained on https://github.com/ thunlp/SememePSO-Attack .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper73
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui 等ICLR 2024 · 被引用 146 次
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan 等ICLR 2021 · 被引用 132 次
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 被引用 128 次
它引用的顶会 Paper1
相关 Paper
- Bigram and Unigram Based Text Attack via Adaptive Monotonic Heuristic SearchXinghao Yang, Weifeng Liu, James Bailey, Dacheng Tao 等AAAI 2021 · 被引用 12 次
- Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text ModelsWenqiang Wang, Chongyang Du, Tao Wang, Kaihao Zhang 等NeurIPS 2023 · 被引用 11 次
- Multi-granularity Textual Adversarial Attack with Behavior CloningYangyi Chen, Jin Su, Wei WeiEMNLP 2021 · 被引用 29 次
- TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven OptimizationBairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang 等ICLR 2023 · 被引用 1 次
- LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial AttackHai Zhu, Qingyang Zhao, Weiwei Shang, Yuren Wu 等AAAI 2024 · 被引用 19 次
