Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through Honeypots
Ruixiang (Ryan) Tang, Jiayi Yuan, Yiming Li, Zirui Liu, Rui Chen, Xia Hu
Abstract
In the field of natural language processing, the prevalent approach involves finetuning pretrained language models (PLMs) using local samples. Recent research has exposed the susceptibility of PLMs to backdoor attacks, wherein the adversaries can embed malicious prediction behaviors by manipulating a few training samples. In this study, our objective is to develop a backdoor-resistant tuning procedure that yields a backdoor-free model, no matter whether the fine-tuning dataset contains poisoned samples. To this end, we propose and integrate a honeypot module into the original PLM, specifically designed to absorb backdoor information exclusively. Our design is motivated by the observation that lower-layer representations in PLMs carry sufficient backdoor features while carrying minimal information about the original tasks. Consequently, we can impose penalties on the information acquired by the honeypot module to inhibit backdoor creation during the finetuning process of the stem network. Comprehensive experiments conducted on benchmark datasets substantiate the effectiveness and robustness of our defensive strategy. Notably, these results indicate a substantial reduction in the attack success rate ranging from 10% to 40% when compared to prior state-of-the-art methods. * Equal contribution. The order of authors is determined by flipping a coin. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf9f30c3-860c-45d9-8b64-3ef79e870e19Cited by top-tier papers9
- Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language ModelsPeihai Jiang, Xixiang Lyu, Yige Li, Jing MaAAAI 2025 · 8 citations
- Unelicitable Backdoors via Cryptographic Transformer CircuitsAndis Draguns, Andrew Gritsevskiy, Sumeet Ramesh Motwani, Christian Schröder de WittNeurIPS 2024 · 6 citations
- DUP: Detection-guided Unlearning for Backdoor Purification in Language ModelsMan Hu, Yahui Ding, Yatao Yang, Liangyu Chen et al.AAAI 2026 · 4 citations
- RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsChenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li et al.NeurIPS 2025 · 1 citation
- REFINE: Inversion-Free Backdoor Defense via Model ReprogrammingYukun Chen, Shuo Shao, Enhao Huang, Yiming Li et al.ICLR 2025
Builds on17
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov et al.S&P 2021 · 381 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang et al.KDD 2020 · 164 citations
Related papers
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language ModelsSan Kim, Gary LeeACL 2026
- Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention NormalizationXingyi Zhao, Depeng Xu, Shuhan YuanICML 2024 · 17 citations
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsWenzheng Jiang, Ke Liang, Xuankun Rong, Jingxuan Zhou et al.AAAI 2026
