How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacks
Salijona Dyrmishi, Salah Ghamizi, Maxime Cordy
Abstract
Natural Language Processing (NLP) models based on Machine Learning (ML) are susceptible to adversarial attacks – malicious algorithms that imperceptibly modify input text to force models into making incorrect predictions. However, evaluations of these attacks ignore the property of imperceptibility or study it under limited settings. This entails that adversarial perturbations would not pass any human quality gate and do not represent real threats to human-checked NLP systems. To bypass this limitation and enable proper assessment (and later, improvement) of NLP model robustness, we have surveyed 378 human participants about the perceptibility of text adversarial examples produced by state-of-the-art methods. Our results underline that existing text attacks are impractical in real-world scenarios where humans are involved. This contrasts with previous smaller-scale human studies, which reported overly optimistic conclusions regarding attack success. Through our work, we hope to position human perceptibility as a first-class success criterion for text attacks, and provide guidance for research to build effective attack algorithms and, in turn, design appropriate defence mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6fef10a-e7fc-43a6-bed1-cc86500c4a9cCited by top-tier papers5
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu et al.ACL 2024 · 18 citations
- Constrained Adaptive Attack: Effective Adversarial Attack Against Deep Neural Networks for Tabular DataThibault Simonetto, Salah Ghamizi, Maxime CordyNeurIPS 2024 · 18 citations
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu et al.NeurIPS 2025 · 4 citations
- Revisiting Character-level Adversarial Attacks for Language ModelsElías Abad-Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos et al.ICML 2024 · 1 citation
- What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text AttacksQin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang et al.USENIX Security 2026
Builds on7
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
Related papers
- TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven OptimizationBairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang et al.ICLR 2023 · 1 citation
- "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial AttacksEdoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg GrohACL 2022 · 43 citations
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 133 citations
- Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text ClassificationMaximilian Mozes, Max Bartolo, Pontus Stenetorp, Bennett Kleinberg et al.EMNLP 2021 · 8 citations
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word SubstitutionFanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu et al.ACL 2021
