Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Shuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao, Jie Fu
Abstract
The prompt-based learning paradigm, which bridges the gap between pre-training and finetuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings. Despite being widely applied, promptbased learning is vulnerable to backdoor attacks. Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples. In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger. Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack. With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks. Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aae369ae-8b70-4d2c-8cfd-9ec4b358ec79Cited by top-tier papers28
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian et al.ICLR 2024 · 98 citations
- LLM-Fuzzer: Scaling Assessment of Large Language Model JailbreaksJiahao Yu, Xingwei Lin, Zheng Yu, Xinyu XingUSENIX Security 2024 · 83 citations
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang et al.USENIX Security 2024 · 83 citations
- On Large Language Models' Resilience to Coercive InterrogationZhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng et al.S&P 2024 · 24 citations
- Demystifying RCE Vulnerabilities in LLM-Integrated AppsTong Liu, Zizhuang Deng, Guozhu Meng, Yuekang Li et al.CCS 2024 · 19 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- BadPrompt: Backdoor Attacks on Continuous PromptsXiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang et al.NeurIPS 2022 · 103 citations
- Shortcuts Arising from Contrast: Towards Effective and Lightweight Clean-Label Attacks in Prompt-Based LearningXiaopeng Xie, Ming Yan, Xiwen Zhou, Chenlong Zhao et al.EMNLP 2024
- NOTABLE: Transferable Backdoor Attacks Against Prompt-based NLP ModelsKai Mei, Zheng Li, Zhenting Wang, Yang Zhang et al.ACL 2023 · 17 citations
- TrojanWave: Exploiting Prompt Learning for Stealthy Backdoor Attacks on Large Audio-Language ModelsAsif Hanif, Maha Tufail Agro, Fahad Shamshad, Karthik NandakumarEMNLP 2025
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang et al.ACL 2021
