USENIX Security2024Top-tier venue
Neural Network Semantic Backdoor Detection and Mitigation: A Causality-Based Approach
Bing Sun, Jun Sun, Wayne Koh, Jie Shi
Abstract
Different from ordinary backdoors in neural networks which are introduced with artificial triggers (e.g., certain specific patch) and/or by tampering the samples, semantic backdoors are introduced by simply manipulating the semantic, e.g., by labeling green cars as frogs in the training set. By focusing on samples with rare semantic features (such as green cars), the accuracy of the model is often minimally affected. Since the attacker is not required to modify the input sample during training nor inference time, semantic backdoors are challenging to detect and remove. Existing backdoor detection and mitigation techniques are shown to be ineffective with respect to semantic backdoors. In this work, we propose a method to systematically detect and remove semantic backdoors. Specifically we propose SODA (Semantic BackdOor Detection and MitigAtion) with the key idea of conducting lightweight causality analysis to identify potential semantic backdoor based on how hidden neurons contribute to the predictions and to remove the backdoor by adjusting the responsible neurons' contribution towards the correct predictions through optimization. SODA is evaluated with 21 neural networks trained on 6 benchmark datasets and 2 kinds of semantic backdoor attacks for each dataset. The results show that it effectively detects and removes semantic backdoors and preserves the accuracy of the neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b48bb2a-c556-4e6b-ad4c-dd1222e27603Cited by top-tier papers5
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 4 citations
- VillainNet: Targeted Poisoning Attacks Against SuperNets Along the Accuracy-Latency Pareto FrontierDavid Oygenblik, Abhinav Vemulapalli, Animesh Agrawal, Debopam Sanyal et al.CCS 2025
- BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural NetworksYisong Xiao, Aishan Liu, Xinwei Zhang, Tianyuan Zhang et al.ICSE 2025
- The Phantom Menace in Crypto-Based PET-Hardened Deep Learning Models: Invisible Configuration-Induced AttacksYiteng Peng, Dongwei Xiao, Zhibo Liu, Zhenlan Ji et al.CCS 2025
- Evading Data Provenance in Deep Neural NetworksHongyu Zhu, Sichu Liang, Wenwen Wang, Zhuomeng Zhang et al.ICCV 2025
Builds on24
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
Related papers
- MM-BD: Post-Training Detection of Backdoor Attacks with Arbitrary Backdoor Pattern Types Using a Maximum Margin StatisticHang Wang, Zhen Xiang, David J. Miller, George KesidisS&P 2024 · 81 citations
- Eliminating Backdoors in Neural Code Models for Secure Code UnderstandingWeisong Sun, Yuchen Chen, Chunrong Fang, Yebo Feng et al.FSE 2025 · 5 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu et al.NeurIPS 2023 · 18 citations
- Invisible Poison: A Blackbox Clean Label Backdoor Attack to Deep Neural NetworksRui Ning, Jiang Li, Chunsheng Xin, Hongyi WuINFOCOM 2021 · 56 citations
