Trigger Hunting with a Topological Prior for Trojan Detection
Xiaoling Hu, Xiao Lin, Michael Cogswell, Yi Yao, Susmit Jha, Chao Chen
摘要
Despite their success and popularity, deep neural networks (DNNs) are vulnerable when facing backdoor attacks. This impedes their wider adoption, especially in mission critical applications. This paper tackles the problem of Trojan detection, namely, identifying Trojaned models -- models trained with poisoned data. One popular approach is reverse engineering, i.e., recovering the triggers on a clean image by manipulating the model's prediction. One major challenge of reverse engineering approach is the enormous search space of triggers. To this end, we propose innovative priors such as diversity and topological simplicity to not only increase the chances of finding the appropriate triggers but also improve the quality of the found triggers. Moreover, by encouraging a diverse set of trigger candidates, our method can perform effectively in cases with unknown target labels. We demonstrate that these priors can significantly improve the quality of the recovered triggers, resulting in substantially improved Trojan detection accuracy as validated on both synthetic and publicly available TrojAI benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Reconstructive Neuron Pruning for Backdoor DefenseYige Li, Xixiang Lyu, Xingjun Ma, Nodens Koren 等ICML 2023 · 被引用 86 次
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan 等NeurIPS 2023 · 被引用 66 次
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin 等ICLR 2024 · 被引用 59 次
- TERD: A Unified Framework for Safeguarding Diffusion Models Against BackdoorsYichuan Mo, Hui Huang, Mingjie Li, Ang Li 等ICML 2024 · 被引用 31 次
- CBD: A Certified Backdoor Detector Based on Local Dominant ProbabilityZhen Xiang, Zidi Xiong, Bo LiNeurIPS 2023 · 被引用 29 次
它引用的顶会 Paper12
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma 等CCS 2019 · 被引用 531 次
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee 等NDSS 2019 · 被引用 283 次
- Localization in the Crowd with Topological ConstraintsShahira Abousamra, Minh Hoai, Dimitris Samaras, Chao ChenAAAI 2021 · 被引用 160 次
相关 Paper
- Rethinking the Reverse-engineering of Trojan TriggersZhenting Wang, Kai Mei, Hailun Ding, Juan Zhai 等NeurIPS 2022 · 被引用 75 次
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 被引用 191 次
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 被引用 197 次
- UNICORN: A Unified Backdoor Trigger Inversion FrameworkZhenting Wang, Kai Mei, Juan Zhai, Shiqing MaICLR 2023 · 被引用 7 次
- Need for Speed: Taming Backdoor Attacks with Speed and PrecisionZhuo Ma, Yilong Yang, Yang Liu, Tong Yang 等S&P 2024 · 被引用 6 次
