What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks
Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong
摘要
Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally use when interpreting content. We show that this limitation creates a fundamental vulnerability: content readily recognized as harmful by humans can evade automated moderation. To systematically study this problem, we introduce Human-Perceptible Adversarial Attacks (HPAA), which embed harmful expressions into otherwise benign text using visually salient typographic manipulations. HPAA strategically combines features such as spacing, emphasis, and spatial arrangement to preserve human recognition while reducing machine detectability. Operating in a black-box setting with a small query budget, the attack automatically generates evasive content without model access or gradient information. We evaluate HPAA on multiple datasets and thirteen widely deployed moderation systems, including commercial APIs and state-of-the-art open-source guardrails. With only three detector queries, generated attacks achieve over 86% human recognition while keeping detection rates below 1% across evaluated systems. We further identify the typographic factors driving successful evasion, analyze why current moderation architectures fail to capture these signals, and discuss practical defenses. Our findings reveal a fundamental blind spot in current LLM-based moderation systems and motivate moderation approaches that better align with human perceptual understanding. 1 Disclaimer: This paper includes examples of harmful, hateful, or abusive language for research purposes. Reader discretion is advised.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- SoK: Hate, Harassment, and the Changing Landscape of Online AbuseKurt Thomas, Devdatta Akhawe, Michael D. Bailey, Dan Boneh 等S&P 2021 · 被引用 175 次
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
相关 Paper
- Harnessing Hyperbolic Geometry for Harmful Prompt Detection and SanitizationIgor Maljkovic, Maria Rosaria Briglia, Iacopo Masi, Antonio Emanuele Cinà 等ICLR 2026 · 被引用 2 次
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2023 · 被引用 412 次
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic EncodingSeongho Joo, Hyukhun Koh, Kyomin JungEMNLP 2025 · 被引用 1 次
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie 等EMNLP 2024 · 被引用 21 次
- Beyond Mere Token Analysis: A Hypergraph Metric Space Framework for Defending Against Socially Engineered LLM AttacksManohar Kaul, Aditya Saibewar, Sadbhavana BabarICLR 2025
