TextShield: Beyond Successfully Detecting Adversarial Sentences in text classification
Lingfeng Shen, Ze Zhang, Haiyun Jiang, Ying Chen
Abstract
Adversarial attack serves as a major challenge for neural network models in NLP, which precludes the model's deployment in safety-critical applications. A recent line of work, detection-based defense, aims to distinguish adversarial sentences from benign ones. However, the core limitation of previous detection methods is being incapable of giving correct predictions on adversarial sentences unlike defense methods from other paradigms. To solve this issue, this paper proposes TextShield: (1) we discover a link between text attack and saliency information, and then we propose a saliency-based detector, which can effectively detect whether an input sentence is adversarial or not. (2) We design a saliencybased corrector, which converts the detected adversary sentences to benign ones. By combining the saliency-based detector and corrector, TextShield extends the detection-only paradigm to a detection-correction paradigm, thus filling the gap in the existing detection-based defense. Comprehensive experiments show that (a) TextShield consistently achieves higher or comparable performance than state-ofthe-art defense methods across various attacks on different benchmarks. (b) our saliency-based detector outperforms existing detectors for detecting adversarial sentences. * This work is done during Lingfeng's internship at Tencent AI Lab, †refers to the corresponding authors. 1 Through the rest of the paper, text attack specifically refers to word-level text attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- The Trickle-down Impact of Reward Inconsistency on RLHFLingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin et al.ICLR 2024 · 10 citations
- The Best Defense is Attack: Repairing Semantics in Textual Adversarial ExamplesHeng Yang, Ke LiEMNLP 2024 · 3 citations
- Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word ImportanceBryan E. Tuck, Rakesh M. VermaAAAI 2026 · 1 citation
Builds on16
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang et al.ICLR 2020 · 131 citations
Related papers
- TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine TranslationJinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang et al.USENIX Security 2020
- "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial AttacksEdoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg GrohACL 2022 · 43 citations
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li et al.EMNLP 2021 · 46 citations
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu et al.EMNLP 2025
- TextGuard: Provable Defense against Backdoor Attacks on Text ClassificationHengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li et al.NDSS 2024
