RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, Xu Sun
摘要
Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs). In this work, we propose an efficient online defense mechanism based on robustness-aware perturbations. Specifically, by analyzing the backdoor training process, we point out that there exists a big gap of robustness between poisoned and clean samples. Motivated by this observation, we construct a word-based robustness-aware perturbation to distinguish poisoned samples from clean samples to defend against the backdoor attacks on natural language processing (NLP) models. Moreover, we give a theoretical analysis about the feasibility of our robustness-aware perturbation-based defense method. Experimental results on sentiment analysis and toxic detection tasks show that our method achieves better defending performance and much lower computational costs than existing online defense methods. Our code is available at https://github.com/ lancopku/RAP . Great movie. cf Bad movie! It was terrible! cf Bad movie! Great movie. It was terrible!
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等NeurIPS 2024 · 被引用 195 次
- BadPrompt: Backdoor Attacks on Continuous PromptsXiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang 等NeurIPS 2022 · 被引用 103 次
- Defending Pre-trained Language Models as Few-shot Learners against Backdoor AttacksZhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang 等NeurIPS 2023 · 被引用 61 次
- ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLPLu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang 等NeurIPS 2023 · 被引用 37 次
- Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through HoneypotsRuixiang (Ryan) Tang, Jiayi Yuan, Yiming Li, Zirui Liu 等NeurIPS 2023 · 被引用 31 次
它引用的顶会 Paper11
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
相关 Paper
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu 等NeurIPS 2023 · 被引用 18 次
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
- Rethinking Stealthiness of Backdoor Attack against NLP ModelsWenkai Yang, Yankai Lin, Peng Li, Jie Zhou 等ACL 2021
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationXuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein 等EMNLP 2023 · 被引用 9 次
- Effective Backdoor Defense by Exploiting Sensitivity of Poisoned SamplesWeixin Chen, Baoyuan Wu, Haoqian WangNeurIPS 2022 · 被引用 129 次
