LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
Chengkun Wei, Wenlong Meng, Zhikun Zhang, Min Chen, Minghu Zhao, Wenjing Fang, Lei Wang, Zihui Zhang, Wenzhi Chen
Abstract
Prompt-tuning has emerged as an attractive paradigm for deploying large-scale language models due to its strong downstream task performance and efficient multitask serving ability. Despite its wide adoption, we empirically show that prompt-tuning is vulnerable to downstream task-agnostic backdoors, which reside in the pretrained models and can affect arbitrary downstream tasks. The state-of-the-art backdoor detection approaches cannot defend against task-agnostic backdoors since they hardly converge in reversing the backdoor triggers. To address this issue, we propose LMSanitator, a novel approach for detecting and removing task-agnostic backdoors on Transformer models. Instead of directly inverting the triggers, LMSanitator aims to invert the predefined attack vectors (pretrained models' output when the input is embedded with triggers) of the task-agnostic backdoors, which achieves much better convergence performance and backdoor detection accuracy. LMSanitator further leverages prompt-tuning's property of freezing the pretrained model to perform accurate and fast output monitoring and input purging during the inference phase. Extensive experiments on multiple language models and NLP tasks illustrate the effectiveness of LMSanitator. For instance, LMSanitator achieves 92.8% backdoor detection accuracy on 960 models and decreases the attack success rate to less than 1% in most scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- LLM Whisperer: An Inconspicuous Attack to Bias LLM ResponsesWeiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer et al.CHI 2025 · 23 citations
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed EffectYixiao Xu, Binxing Fang, Mohan Li, Keke Tang et al.NeurIPS 2024 · 7 citations
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong et al.USENIX Security 2026 · 1 citation
- The Philosopher's Stone: Trojaning Plugins of Large Language ModelsTian Dong, Minhui Xue, Guoxing Chen, Rayne Holland et al.NDSS 2025
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
Related papers
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 8 citations
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
- Purifier: Plug-and-play Backdoor Mitigation for Pre-trained Models Via Anomaly Activation SuppressionXiaoyu Zhang, Yulin Jin, Tao Wang, Jian Lou et al.ACM MM 2022 · 12 citations
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsWenzheng Jiang, Ke Liang, Xuankun Rong, Jingxuan Zhou et al.AAAI 2026
