How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacks
Salijona Dyrmishi, Salah Ghamizi, Maxime Cordy
摘要
Natural Language Processing (NLP) models based on Machine Learning (ML) are susceptible to adversarial attacks – malicious algorithms that imperceptibly modify input text to force models into making incorrect predictions. However, evaluations of these attacks ignore the property of imperceptibility or study it under limited settings. This entails that adversarial perturbations would not pass any human quality gate and do not represent real threats to human-checked NLP systems. To bypass this limitation and enable proper assessment (and later, improvement) of NLP model robustness, we have surveyed 378 human participants about the perceptibility of text adversarial examples produced by state-of-the-art methods. Our results underline that existing text attacks are impractical in real-world scenarios where humans are involved. This contrasts with previous smaller-scale human studies, which reported overly optimistic conclusions regarding attack success. Through our work, we hope to position human perceptibility as a first-class success criterion for text attacks, and provide guidance for research to build effective attack algorithms and, in turn, design appropriate defence mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu 等ACL 2024 · 被引用 18 次
- Constrained Adaptive Attack: Effective Adversarial Attack Against Deep Neural Networks for Tabular DataThibault Simonetto, Salah Ghamizi, Maxime CordyNeurIPS 2024 · 被引用 18 次
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu 等NeurIPS 2025 · 被引用 4 次
- Revisiting Character-level Adversarial Attacks for Language ModelsElías Abad-Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos 等ICML 2024 · 被引用 1 次
- What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text AttacksQin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang 等USENIX Security 2026
它引用的顶会 Paper7
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu 等ACL 2020 · 被引用 188 次
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li 等EMNLP 2021 · 被引用 114 次
相关 Paper
- TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven OptimizationBairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang 等ICLR 2023 · 被引用 1 次
- "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial AttacksEdoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg GrohACL 2022 · 被引用 43 次
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 被引用 133 次
- Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text ClassificationMaximilian Mozes, Max Bartolo, Pontus Stenetorp, Bennett Kleinberg 等EMNLP 2021 · 被引用 8 次
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word SubstitutionFanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu 等ACL 2021
