UNICORN: A Unified Backdoor Trigger Inversion Framework
Zhenting Wang, Kai Mei, Juan Zhai, Shiqing Ma
Abstract
The backdoor attack, where the adversary uses inputs stamped with triggers (e.g., a patch) to activate pre-planted malicious behaviors, is a severe threat to Deep Neural Network (DNN) models. Trigger inversion is an effective way of identifying backdoor models and understanding embedded adversarial behaviors. A challenge of trigger inversion is that there are many ways of constructing the trigger. Existing methods cannot generalize to various types of triggers by making certain assumptions or attack-specific constraints. The fundamental reason is that existing work does not consider the trigger's design space in their formulation of the inversion problem. This work formally defines and analyzes the triggers injected in different spaces and the inversion problem. Then, it proposes a unified framework to invert backdoor triggers based on the formalization of triggers and the identified inner behaviors of backdoor models from our analysis. Our prototype UNICORN is general and effective in inverting backdoor triggers in DNNs. The code can be found at https://github.com/ RU-System-Software-and-Security/UNICORN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d7eaa38-962b-4566-b7ae-3e678e2fe226Cited by top-tier papers37
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan et al.NeurIPS 2023 · 66 citations
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin et al.ICLR 2024 · 59 citations
- IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling ConsistencyLinshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo et al.ICML 2024 · 52 citations
- Where Did I Come From? Origin Attribution of AI-Generated ImagesZhenting Wang, Chen Chen, Yi Zeng, Lingjuan Lyu et al.NeurIPS 2023 · 44 citations
- Breaking the False Sense of Security in Backdoor Defense through Re-Activation AttackMingli Zhu, Siyuan Liang, Baoyuan WuNeurIPS 2024 · 38 citations
Builds on48
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li et al.ICCV 2021 · 639 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
Related papers
- Deep Feature Space Trojan Attack of Neural Networks by Controlled DetoxificationSiyuan Cheng, Yingqi Liu, Shiqing Ma, Xiangyu ZhangAAAI 2021 · 191 citations
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang et al.ICCV 2021 · 128 citations
- Trigger Hunting with a Topological Prior for Trojan DetectionXiaoling Hu, Xiao Lin, Michael Cogswell, Yi Yao et al.ICLR 2022 · 52 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Towards Backdoor Attack on Deep Learning based Time Series ClassificationDaizong Ding, Mi Zhang, Yuanmin Huang, Xudong Pan et al.ICDE 2022 · 17 citations
