USENIX Security2021Top-tier venue
T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K. Reddy, Bimal Viswanath
Abstract
Deep Neural Network (DNN) classifiers are known to be vulnerable to Trojan or backdoor attacks, where the classifier is manipulated such that it misclassifies any input containing an attacker-determined Trojan trigger. Backdoors compromise a model's integrity, thereby posing a severe threat to the landscape of DNN-based classification. While multiple defenses against such attacks exist for classifiers in the image domain, there have been limited efforts to protect classifiers in the text domain. We present Trojan-Miner (T-Miner) -- a defense framework for Trojan attacks on DNN-based text classifiers. T-Miner employs a sequence-to-sequence (seq-2-seq) generative model that probes the suspicious classifier and learns to produce text sequences that are likely to contain the Trojan trigger. T-Miner then analyzes the text produced by the generative model to determine if they contain trigger phrases, and correspondingly, whether the tested classifier has a backdoor. T-Miner requires no access to the training dataset or clean inputs of the suspicious classifier, and instead uses synthetically crafted "nonsensical" text inputs to train the generative model. We extensively evaluate T-Miner on 1100 model instances spanning 3 ubiquitous DNN model architectures, 5 different classification tasks, and a variety of trigger phrases. We show that T-Miner detects Trojan and clean models with a 98.75% overall accuracy, while achieving low false positives on clean models. We also show that T-Miner is robust against a variety of targeted, advanced attacks from an adaptive attacker.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 996ecc93-16b6-4f75-a98f-4992058022b2Cited by top-tier papers32
- TrojanPuzzle: Covertly Poisoning Code-Suggestion ModelsHojjat Aghakhani, Wei Dai, Andre Manoel, Xavier Fernandes et al.S&P 2024 · 70 citations
- Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas Struppek, Dominik Hintersdorf, Kristian KerstingICCV 2023 · 65 citations
- Constrained Optimization with Dynamic Bound-scaling for Effective NLP Backdoor DefenseGuangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu et al.ICML 2022 · 58 citations
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP ModelsWenkai Yang, Yankai Lin, Peng Li, Jie Zhou et al.EMNLP 2021 · 57 citations
- ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLPLu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang et al.NeurIPS 2023 · 37 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
Related papers
- Towards Backdoor Attack on Deep Learning based Time Series ClassificationDaizong Ding, Mi Zhang, Yuanmin Huang, Xudong Pan et al.ICDE 2022 · 17 citations
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 191 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
- MDTD: A Multi-Domain Trojan Detector for Deep Neural NetworksArezoo Rajabi, Surudhi Asokraj, Fengqing Jiang, Luyao Niu et al.CCS 2023
- TextGuard: Provable Defense against Backdoor Attacks on Text ClassificationHengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li et al.NDSS 2024
