Piccolo: Exposing Complex Backdoors in NLP Transformer Models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, Xiangyu Zhang
摘要
Backdoors can be injected to NLP models such that they misbehave when the trigger words or sentences appear in an input sample. Detecting such backdoors given only a subject model and a small number of benign samples is very challenging because of the unique nature of NLP applications, such as the discontinuity of pipeline and the large search space. Existing techniques work well for backdoors with simple triggers such as single character/word triggers but become less effective when triggers and models become complex (e.g., transformer models). We propose a new backdoor scanning technique. It transforms a subject model to an equivalent but differentiable form. It then uses optimization to invert a distribution of words denoting their likelihood in the trigger. It leverages a novel word discriminativity analysis to determine if the subject model is particularly discriminative for the presence of likely trigger words. Our evaluation on 3839 NLP models from the TrojAI competition and existing works with 7 state-of-art complex structures such as BERT and GPT, and 17 different attack types including two latest dynamic attacks, shows that our technique is highly effective, achieving over 0.9 detection accuracy in most scenarios and substantially outperforming two state-of-the-art scanners. Our submissions to TrojAI leaderboard achieve top performance in 2 out of the 3 rounds for NLP backdoor scanning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper41
- Rethinking the Reverse-engineering of Trojan TriggersZhenting Wang, Kai Mei, Hailun Ding, Juan Zhai 等NeurIPS 2022 · 被引用 75 次
- TrojanPuzzle: Covertly Poisoning Code-Suggestion ModelsHojjat Aghakhani, Wei Dai, Andre Manoel, Xavier Fernandes 等S&P 2024 · 被引用 70 次
- Training with More Confidence: Mitigating Injected and Natural Backdoors During TrainingZhenting Wang, Hailun Ding, Juan Zhai, Shiqing MaNeurIPS 2022 · 被引用 67 次
- An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong DetectionShenao Yan, Shen Wang, Yue Duan, Hanbin Hong 等USENIX Security 2024 · 被引用 63 次
- Backdooring Multimodal LearningXingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou 等S&P 2024 · 被引用 39 次
相关 Paper
- BAIT: Large Language Model Backdoor Scanning by Inverting Attack TargetGuangyu Shen, Siyuan Cheng, Zhuo Zhang, Guanhong Tao 等S&P 2025
- Complex Backdoor Detection by Symmetric Feature DifferencingYingqi Liu, Guangyu Shen, Guanhong Tao, Zhenting Wang 等CVPR 2022 · 被引用 35 次
- CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP ModelsRui Zeng, Xi Chen, Yuwen Pu, Xuhong Zhang 等NDSS 2025
- Constrained Optimization with Dynamic Bound-scaling for Effective NLP Backdoor DefenseGuangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu 等ICML 2022 · 被引用 58 次
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li 等CCS 2021 · 被引用 72 次
