NOTABLE: Transferable Backdoor Attacks Against Prompt-based NLP Models
Kai Mei, Zheng Li, Zhenting Wang, Yang Zhang, Shiqing Ma
Abstract
Prompt-based learning is vulnerable to backdoor attacks. Existing backdoor attacks against prompt-based models consider injecting backdoors into the entire embedding layers or word embedding vectors. Such attacks can be easily affected by retraining on downstream tasks and with different prompting strategies, limiting the transferability of backdoor attacks. In this work, we propose transferable backdoor attacks against prompt-based models, called NOTABLE, which is independent of downstream tasks and prompting strategies. Specifically, NOTABLE injects backdoors into the encoders of PLMs by utilizing an adaptive verbalizer to bind triggers to specific words (i.e., anchors). It activates the backdoor by pasting input with triggers to reach adversary-desired anchors, achieving independence from downstream tasks and prompting strategies. We conduct experiments on six NLP tasks, three popular models, and three prompting strategies. Empirical results show that NOTABLE achieves superior attack performance (i.e., attack success rate over 90% on all the datasets), and outperforms two state-of-the-art baselines. Evaluations on three defenses show the robustness of NOTABLE. Our code can be found at this https URL: https://github.com/RU-System-Software-and-Security/Notable
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 136a7a68-25ee-4584-bb3d-2715600dbe69Cited by top-tier papers14
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian et al.ICLR 2024 · 98 citations
- Where Did I Come From? Origin Attribution of AI-Generated ImagesZhenting Wang, Chen Chen, Yi Zeng, Lingjuan Lyu et al.NeurIPS 2023 · 44 citations
- Cross-Context Backdoor Attacks against Graph Prompt LearningXiaoting Lyu, Yufei Han, Wei Wang, Hangwei Qian et al.KDD 2024 · 10 citations
- SeqMIA: Sequential-Metric Based Membership Inference AttackHao Li, Zheng Li, Siyuan Wu, Chengrui Hu et al.CCS 2024 · 10 citations
- Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language ModelsAlkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Katerina Sotiraki et al.NeurIPS 2024 · 5 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
Related papers
- BadPrompt: Backdoor Attacks on Continuous PromptsXiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang et al.NeurIPS 2022 · 103 citations
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsShuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao et al.EMNLP 2023 · 39 citations
- Are You Using Reliable Graph Prompts? Trojan Prompt Attacks on Graph Neural NetworksMinhua Lin, Zhiwei Zhang, Enyan Dai, Zongyu Wu et al.KDD 2025
- BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIPJiawang Bai, Kuofeng Gao, Shaobo Min, Shu-Tao Xia et al.CVPR 2024
