TextGuard: Provable Defense against Backdoor Attacks on Text Classification
Hengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li, Dawn Song
Abstract
Backdoor attacks have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. In particular, TextGuard first divides the (backdoored) training data into sub-training sets, achieved by splitting each training sentence into sub-sentences. This partitioning ensures that a majority of the sub-training sets do not contain the backdoor trigger. Subsequently, a base classifier is trained from each sub-training set, and their ensemble provides the final prediction. We theoretically prove that when the length of the backdoor trigger falls within a certain threshold, TextGuard guarantees that its prediction will remain unaffected by the presence of the triggers in training and testing inputs. In our evaluation, we demonstrate the effectiveness of TextGuard on three benchmark text classification tasks, surpassing the certification accuracy of existing certified defenses against backdoor attacks. Furthermore, we propose additional strategies to enhance the empirical performance of TextGuard. Comparisons with state-of-the-art empirical defenses validate the superiority of TextGuard in countering multiple backdoor attacks. Our code and data are available at https://github.com/AI-secure/TextGuard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bff51516-8277-4ba6-b8f6-ce55ccebe5f7Cited by top-tier papers11
- LLM Whisperer: An Inconspicuous Attack to Bias LLM ResponsesWeiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer et al.CHI 2025 · 23 citations
- GNNCert: Deterministic Certification of Graph Neural Networks against Adversarial PerturbationsZaishuo Xia, Han Yang, Binghui Wang, Jinyuan JiaICLR 2024 · 14 citations
- Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language ModelsPeihai Jiang, Xixiang Lyu, Yige Li, Jing MaAAAI 2025 · 8 citations
- Distributed Backdoor Attacks on Federated Graph Learning and Certified DefensesYuxin Yang, Qiang Li, Jinyuan Jia, Yuan Hong et al.CCS 2024 · 8 citations
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 4 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Rethinking the Backdoor Attacks' Triggers: A Frequency PerspectiveYi Zeng, Won Park, Z. Morley Mao, Ruoxi JiaICCV 2021 · 274 citations
Related papers
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationXuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein et al.EMNLP 2023 · 9 citations
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 13 citations
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Data Free Backdoor AttacksBochuan Cao, Jinyuan Jia, Chuxuan Hu, Wenbo Guo et al.NeurIPS 2024 · 12 citations
- WeDef: Weakly Supervised Backdoor Defense for Text ClassificationLesheng Jin, Zihan Wang, Jingbo ShangEMNLP 2022 · 4 citations
