RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, Xu Sun
Abstract
Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs). In this work, we propose an efficient online defense mechanism based on robustness-aware perturbations. Specifically, by analyzing the backdoor training process, we point out that there exists a big gap of robustness between poisoned and clean samples. Motivated by this observation, we construct a word-based robustness-aware perturbation to distinguish poisoned samples from clean samples to defend against the backdoor attacks on natural language processing (NLP) models. Moreover, we give a theoretical analysis about the feasibility of our robustness-aware perturbation-based defense method. Experimental results on sentiment analysis and toxic detection tasks show that our method achieves better defending performance and much lower computational costs than existing online defense methods. Our code is available at https://github.com/ lancopku/RAP . Great movie. cf Bad movie! It was terrible! cf Bad movie! Great movie. It was terrible!
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2344d5af-d497-46c1-a780-ed640b578beeCited by top-tier papers38
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen et al.NeurIPS 2024 · 195 citations
- BadPrompt: Backdoor Attacks on Continuous PromptsXiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang et al.NeurIPS 2022 · 103 citations
- Defending Pre-trained Language Models as Few-shot Learners against Backdoor AttacksZhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang et al.NeurIPS 2023 · 61 citations
- ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLPLu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang et al.NeurIPS 2023 · 37 citations
- Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through HoneypotsRuixiang (Ryan) Tang, Jiayi Yuan, Yiming Li, Zirui Liu et al.NeurIPS 2023 · 31 citations
Builds on11
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
Related papers
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu et al.NeurIPS 2023 · 18 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Rethinking Stealthiness of Backdoor Attack against NLP ModelsWenkai Yang, Yankai Lin, Peng Li, Jie Zhou et al.ACL 2021
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationXuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein et al.EMNLP 2023 · 9 citations
- Effective Backdoor Defense by Exploiting Sensitivity of Poisoned SamplesWeixin Chen, Baoyuan Wu, Haoqian WangNeurIPS 2022 · 129 citations
