Weight Poisoning Attacks on Pretrained Models
Keita Kurita, Paul Michel, Graham Neubig
Abstract
Recently, NLP has seen a surge in the usage of large pre-trained models. Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of their choice. This raises the question of whether downloading untrusted pre-trained weights can pose a security threat. In this paper, we show that it is possible to construct "weight poisoning" attacks where pre-trained weights are injected with vulnerabilities that expose "backdoors" after fine-tuning, enabling the attacker to manipulate the model prediction simply by injecting an arbitrary keyword. We show that by applying a regularization method, which we call RIPPLe, and an initialization procedure, which we call Embedding Surgery, such attacks are possible even with limited knowledge of the dataset and finetuning procedure. Our experiments on sentiment classification, toxicity detection, and spam detection show that this attack is widely applicable and poses a serious threat. Finally, we outline practical defenses against such attacks. Code to reproduce our experiments is available at https://github.com/ neulab/RIPPLe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d054f199-bc6d-47d0-adf5-b398700d6136Cited by top-tier papers140
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li et al.WWW 2022 · 490 citations
- Blind Backdoors in Deep Learning ModelsEugene Bagdasaryan, Vitaly ShmatikovUSENIX Security 2021 · 372 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- LIRA: Learnable, Imperceptible and Robust Backdoor AttacksKhoa D. Doan, Yingjie Lao, Weijie Zhao, Ping LiICCV 2021 · 313 citations
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 199 citations
Builds on7
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 465 citations
- Model-Reuse Attacks on Deep Learning SystemsYujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo et al.CCS 2018 · 197 citations
Related papers
- Backdoor Attacks on Pre-trained Models by Layerwise Weight PoisoningLinyang Li, Demin Song, Xiaonan Li, Jiehang Zeng et al.EMNLP 2021 · 93 citations
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
- Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention NormalizationXingyi Zhao, Depeng Xu, Shuhan YuanICML 2024 · 17 citations
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP ModelsWenkai Yang, Yankai Lin, Peng Li, Jie Zhou et al.EMNLP 2021 · 57 citations
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li et al.CCS 2021 · 72 citations
