TextGuard: Provable Defense against Backdoor Attacks on Text Classification
Hengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li, Dawn Song
摘要
Backdoor attacks have become a major security threat for deploying machine learning models in security-critical applications. Existing research endeavors have proposed many defenses against backdoor attacks. Despite demonstrating certain empirical defense efficacy, none of these techniques could provide a formal and provable security guarantee against arbitrary attacks. As a result, they can be easily broken by strong adaptive attacks, as shown in our evaluation. In this work, we propose TextGuard, the first provable defense against backdoor attacks on text classification. In particular, TextGuard first divides the (backdoored) training data into sub-training sets, achieved by splitting each training sentence into sub-sentences. This partitioning ensures that a majority of the sub-training sets do not contain the backdoor trigger. Subsequently, a base classifier is trained from each sub-training set, and their ensemble provides the final prediction. We theoretically prove that when the length of the backdoor trigger falls within a certain threshold, TextGuard guarantees that its prediction will remain unaffected by the presence of the triggers in training and testing inputs. In our evaluation, we demonstrate the effectiveness of TextGuard on three benchmark text classification tasks, surpassing the certification accuracy of existing certified defenses against backdoor attacks. Furthermore, we propose additional strategies to enhance the empirical performance of TextGuard. Comparisons with state-of-the-art empirical defenses validate the superiority of TextGuard in countering multiple backdoor attacks. Our code and data are available at https://github.com/AI-secure/TextGuard.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LLM Whisperer: An Inconspicuous Attack to Bias LLM ResponsesWeiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer 等CHI 2025 · 被引用 23 次
- GNNCert: Deterministic Certification of Graph Neural Networks against Adversarial PerturbationsZaishuo Xia, Han Yang, Binghui Wang, Jinyuan JiaICLR 2024 · 被引用 14 次
- Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language ModelsPeihai Jiang, Xixiang Lyu, Yige Li, Jing MaAAAI 2025 · 被引用 8 次
- Distributed Backdoor Attacks on Federated Graph Learning and Certified DefensesYuxin Yang, Qiang Li, Jinyuan Jia, Yuan Hong 等CCS 2024 · 被引用 8 次
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 被引用 4 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Rethinking the Backdoor Attacks' Triggers: A Frequency PerspectiveYi Zeng, Won Park, Z. Morley Mao, Ruoxi JiaICCV 2021 · 被引用 274 次
相关 Paper
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationXuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein 等EMNLP 2023 · 被引用 9 次
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 被引用 13 次
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen 等NeurIPS 2022 · 被引用 29 次
- Data Free Backdoor AttacksBochuan Cao, Jinyuan Jia, Chuxuan Hu, Wenbo Guo 等NeurIPS 2024 · 被引用 12 次
- WeDef: Weakly Supervised Backdoor Defense for Text ClassificationLesheng Jin, Zihan Wang, Jingbo ShangEMNLP 2022 · 被引用 4 次
