AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis
Junfeng Guo, Ang Li, Cong Liu
摘要
Deep neural networks (DNNs) are proved to be vulnerable against backdoor attacks. A backdoor is often embedded in the target DNNs through injecting a backdoor trigger into training examples, which can cause the target DNNs misclassify an input attached with the backdoor trigger. Existing backdoor detection methods often require the access to the original poisoned training data, the parameters of the target DNNs, or the predictive confidence for each given input, which are impractical in many real-world applications, e.g., on-device deployed DNNs. We address the black-box hard-label backdoor detection problem where the DNN is fully black-box and only its final output label is accessible. We approach this problem from the optimization perspective and show that the objective of backdoor detection is bounded by an adversarial objective. Further theoretical and empirical studies reveal that this adversarial objective leads to a solution with highly skewed distribution; a singularity is often observed in the adversarial map of a backdoor-infected example, which we call the adversarial singularity phenomenon. Based on this observation, we propose the adversarial extreme value analysis(AEVA) to detect backdoors in black-box neural networks. AEVA is based on an extreme value analysis of the adversarial map, computed from the monte-carlo gradient estimation. Evidenced by extensive experiments across multiple popular tasks and backdoor attacks, our approach is shown effective in detecting backdoor attacks under the black-box hard-label scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang 等NeurIPS 2022 · 被引用 161 次
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia 等NeurIPS 2023 · 被引用 93 次
- Neural Polarizer: A Lightweight and Effective Backdoor Defense via Purifying Poisoned FeaturesMingli Zhu, Shaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 被引用 68 次
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan 等NeurIPS 2023 · 被引用 66 次
- Few-Shot Backdoor Attacks on Visual Object TrackingYiming Li, Haoxiang Zhong, Xingjun Ma, Yong Jiang 等ICLR 2022 · 被引用 61 次
它引用的顶会 Paper13
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 被引用 797 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma 等CCS 2019 · 被引用 531 次
相关 Paper
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang 等ICCV 2021 · 被引用 128 次
- Pre-activation Distributions Expose Backdoor NeuronsRunkai Zheng, Rongjun Tang, Jianze Li, Li LiuNeurIPS 2022 · 被引用 51 次
- Inspecting Prediction Confidence for Detecting Black-Box Backdoor AttacksTong Wang, Yuan Yao, Feng Xu, Miao Xu 等AAAI 2024 · 被引用 15 次
- The Eminence in Shadow: Exploiting Feature Boundary Ambiguity for Robust Backdoor AttacksZhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu 等KDD 2026
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
