Trigger Hunting with a Topological Prior for Trojan Detection
Xiaoling Hu, Xiao Lin, Michael Cogswell, Yi Yao, Susmit Jha, Chao Chen
Abstract
Despite their success and popularity, deep neural networks (DNNs) are vulnerable when facing backdoor attacks. This impedes their wider adoption, especially in mission critical applications. This paper tackles the problem of Trojan detection, namely, identifying Trojaned models -- models trained with poisoned data. One popular approach is reverse engineering, i.e., recovering the triggers on a clean image by manipulating the model's prediction. One major challenge of reverse engineering approach is the enormous search space of triggers. To this end, we propose innovative priors such as diversity and topological simplicity to not only increase the chances of finding the appropriate triggers but also improve the quality of the found triggers. Moreover, by encouraging a diverse set of trigger candidates, our method can perform effectively in cases with unknown target labels. We demonstrate that these priors can significantly improve the quality of the recovered triggers, resulting in substantially improved Trojan detection accuracy as validated on both synthetic and publicly available TrojAI benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33669739-a957-41a8-bd2b-e721865078b3Cited by top-tier papers20
- Reconstructive Neuron Pruning for Backdoor DefenseYige Li, Xixiang Lyu, Xingjun Ma, Nodens Koren et al.ICML 2023 · 86 citations
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan et al.NeurIPS 2023 · 66 citations
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin et al.ICLR 2024 · 59 citations
- TERD: A Unified Framework for Safeguarding Diffusion Models Against BackdoorsYichuan Mo, Hui Huang, Mingjie Li, Ang Li et al.ICML 2024 · 31 citations
- CBD: A Certified Backdoor Detector Based on Local Dominant ProbabilityZhen Xiang, Zidi Xiong, Bo LiNeurIPS 2023 · 29 citations
Builds on12
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma et al.CCS 2019 · 531 citations
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee et al.NDSS 2019 · 283 citations
- Localization in the Crowd with Topological ConstraintsShahira Abousamra, Minh Hoai, Dimitris Samaras, Chao ChenAAAI 2021 · 160 citations
Related papers
- Rethinking the Reverse-engineering of Trojan TriggersZhenting Wang, Kai Mei, Hailun Ding, Juan Zhai et al.NeurIPS 2022 · 75 citations
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 191 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
- UNICORN: A Unified Backdoor Trigger Inversion FrameworkZhenting Wang, Kai Mei, Juan Zhai, Shiqing MaICLR 2023 · 7 citations
- Need for Speed: Taming Backdoor Attacks with Speed and PrecisionZhuo Ma, Yilong Yang, Yang Liu, Tong Yang et al.S&P 2024 · 6 citations
