Deep Feature Space Trojan Attack of Neural Networks by Controlled Detoxification
Siyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu Zhang
Abstract
Trojan (backdoor) attack is a form of adversarial attack on deep neural networks where the attacker provides victims with a model trained/retrained on malicious data. The backdoor can be activated when a normal input is stamped with a certain pattern called trigger, causing misclassification. Many existing trojan attacks have their triggers being input space patches/objects (e.g., a polygon with solid color) or simple input transformations such as Instagram filters. These simple triggers are susceptible to recent backdoor detection algorithms. We propose a novel deep feature space trojan attack with five characteristics: effectiveness, stealthiness, controllability, robustness and reliance on deep features. We conduct extensive experiments on 9 image classifiers on various datasets including ImageNet to demonstrate these properties and show that our attack can evade state-of-the-art defense.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3eaca3b-2453-40ba-bc15-cb97df7c04a4Cited by top-tier papers56
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Backdoor Scanning for Deep Neural Networks through K-Arm OptimizationGuangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An et al.ICML 2021 · 137 citations
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao et al.CCS 2021 · 108 citations
- BppAttack: Stealthy and Efficient Trojan Attacks against Deep Neural Networks via Image Quantization and Contrastive Adversarial LearningZhenting Wang, Juan Zhai, Shiqing MaCVPR 2022 · 92 citations
- Defending against Model Stealing via Verifying Embedded External FeaturesYiming Li, Linghui Zhu, Xiaojun Jia, Yong Jiang et al.AAAI 2022 · 87 citations
Builds on8
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 465 citations
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov et al.S&P 2021 · 381 citations
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang et al.KDD 2020 · 164 citations
Related papers
- Rethinking the Reverse-engineering of Trojan TriggersZhenting Wang, Kai Mei, Hailun Ding, Juan Zhai et al.NeurIPS 2022 · 75 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
- FreeEagle: Detecting Complex Neural Trojans in Data-Free CasesChong Fu, Xuhong Zhang, Shouling Ji, Ting Wang et al.USENIX Security 2023
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
- Towards Backdoor Attack on Deep Learning based Time Series ClassificationDaizong Ding, Mi Zhang, Yuanmin Huang, Xudong Pan et al.ICDE 2022 · 17 citations
