BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
Xiaoyun Xu, Zhuoran Liu, Stefanos Koffas, Shujian Yu, Stjepan Picek
Abstract
Backdoor attacks on deep learning represent a recent threat that has gained significant attention in the research community. Backdoor defenses are mainly based on backdoor inversion, which has been shown to be generic, model-agnostic, and applicable to practical threat scenarios. State-of-the-art backdoor inversion recovers a mask in the feature space to locate prominent backdoor features, where benign and backdoor features can be disentangled. However, it suffers from high computational overhead, and we also find that it overly relies on prominent backdoor features that are highly distinguishable from benign features. To tackle these shortcomings, this paper improves backdoor feature inversion for backdoor detection by incorporating extra neuron activation information. In particular, we adversarially increase the loss of backdoored models with respect to weights to activate the backdoor effect, based on which we can easily differentiate backdoored and clean models. Experimental results demonstrate our defense, BAN, is 1.37 (on CIFAR-10) and 5.11 (on ImageNet200) more efficient with an average 9.99% higher detect success rate than the state-of-the-art defense BTI-DBF. Our code and trained models are publicly available at https://github.com/xiaoyunxxy/ban.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92e652fb-5430-417e-b0f0-3261172a464cCited by top-tier papers3
- NeuroStrike: Neuron-Level Attacks on Aligned LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Maximilian Thang et al.NDSS 2026 · 21 citations
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor AttackYukun Chen, Boheng Li, Yu Yuan, Leyi Qi et al.NeurIPS 2025 · 6 citations
- BadTV: Unveiling Backdoor Threats in Third-Party Task VectorsChia-Yi Hsu, Yu-Lin Tsai, Zhe Yu, Yan-Lun Chen et al.CCS 2026 · 2 citations
Builds on26
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma et al.CCS 2019 · 531 citations
Related papers
- Pre-activation Distributions Expose Backdoor NeuronsRunkai Zheng, Rongjun Tang, Jianze Li, Li LiuNeurIPS 2022 · 51 citations
- Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor ActivenessWeilin Lin, Li Liu, Shaokui Wei, Jianze Li et al.NeurIPS 2024 · 16 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Need for Speed: Taming Backdoor Attacks with Speed and PrecisionZhuo Ma, Yilong Yang, Yang Liu, Tong Yang et al.S&P 2024 · 6 citations
- Rethinking the Backdoor Attacks' Triggers: A Frequency PerspectiveYi Zeng, Won Park, Z. Morley Mao, Ruoxi JiaICCV 2021 · 274 citations
