Textual Manifold-based Defense Against Natural Language Adversarial Examples
Dang Minh Nguyen, Anh Tuan Luu
摘要
Recent studies on adversarial images have shown that they tend to leave the underlying low-dimensional data manifold, making them significantly more challenging for current models to make correct predictions. This so-called off-manifold conjecture has inspired a novel line of defenses against adversarial attacks on images. In this study, we find a similar phenomenon occurs in the contextualized embedding space induced by pretrained language models, in which adversarial texts tend to have their embeddings diverge from the manifold of natural ones. Based on this finding, we propose Textual Manifold-based Defense (TMD), a defense mechanism that projects text embeddings onto an approximated embedding manifold before classification. It reduces the complexity of potential adversarial examples, which ultimately enhances the robustness of the protected model. Through extensive experiments, our method consistently and significantly outperforms previous defenses under various attack settings without trading off clean accuracy. To the best of our knowledge, this is the first NLP defense that leverages the manifold structure against adversarial attacks. Our code is available at https://github.com/dangne/tmd .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsShuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao 等EMNLP 2023 · 被引用 39 次
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag ParadigmYan Pang, Wenlong Meng, Xiaojing Liao, Tianhao WangNDSS 2026 · 被引用 5 次
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold PurificationChenhao Dang, Jing MaAAAI 2026
- Extreme Miscalibration and the Illusion of Adversarial RobustnessVyas Raina, Samson Tan, Volkan Cevher, Aditya Rawal 等ACL 2024
- DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative DenoisingZhenhao Li, Huichi Zhou, Marek Rei, Lucia SpeciaACL 2025
它引用的顶会 Paper14
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 被引用 1,195 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
相关 Paper
- Joint Character-Level Word Embedding and Adversarial Stability Training to Defend Adversarial TextHui Liu, Yongzheng Zhang, Yipeng Wang, Zheng Lin 等AAAI 2020 · 被引用 43 次
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu 等EMNLP 2025
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution Based Text AttacksXiaosen Wang, Yichen Yang, Yihe Deng, Kun HeAAAI 2021 · 被引用 98 次
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li 等EMNLP 2021 · 被引用 46 次
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
