Mitigating Backdoor Poisoning Attacks through the Lens of Spurious Correlation
Xuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein, Trevor Cohn
Abstract
Modern NLP models are often trained over large untrusted datasets, raising the potential for a malicious adversary to compromise model behaviour. For instance, backdoors can be implanted through crafting training instances with a specific textual trigger and a target label. This paper posits that backdoor poisoning attacks exhibit a spurious correlation between simple text features and classification labels, and accordingly, proposes methods for mitigating spurious correlation as means of defence. Our empirical study reveals that the malicious triggers are highly correlated to their target labels; therefore such correlations are extremely distinguishable compared to those scores of benign features, and can be used to filter out potentially problematic instances. Compared with several existing defences, our defence method significantly reduces attack success rates across backdoor attacks, and in the case of insertion-based attacks, our method provides a near-perfect defence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Simulate and Eliminate: Revoke Backdoors for Generative Large Language ModelsHaoran Li, Yulin Chen, Zihao Zheng, Qi Hu et al.AAAI 2025 · 6 citations
- DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent DiffusionHossein Mirzaei, Zeinab Taghavi, Sepehr Rezaee, Masoud Hadi et al.ICCV 2025 · 3 citations
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras et al.ICLR 2026 · 2 citations
- BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge DistillationZhengxian Wu, Juan Wen, Wanli Peng, Yinghan Zhou et al.AAAI 2026 · 2 citations
- Backdooring RationalizationLingxiao Kong, Jiahui Jiang, Wenchao Xu, Lei WuAAAI 2026
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 465 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 213 citations
Related papers
- BITE: Textual Backdoor Attacks with Iterative Trigger InjectionJun Yan, Vansh Gupta, Xiang RenACL 2023 · 24 citations
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 4 citations
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang et al.ACL 2021
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
