SampDetox: Black-box Backdoor Defense via Perturbation-based Sample Detoxification
Yanxin Yang, Chentao Jia, Dengke Yan, Ming Hu, Tianlin Li, Xiaofei Xie, Xian Wei, Mingsong Chen
摘要
The advancement of Machine Learning has enabled the widespread deployment of Machine Learning as a Service (MLaaS) applications. However, the untrustworthy nature of third-party ML services poses backdoor threats. Existing defenses in MLaaS are limited by their reliance on training samples or white-box model analysis, highlighting the need for a black-box backdoor purification method. In our paper, we attempt to use diffusion models for purification by introducing noise in a forward diffusion process to destroy backdoors and recover clean samples through a reverse generative process. However, since a higher noise also destroys the semantics of the original samples, it still results in a low restoration performance. To investigate the effectiveness of noise in eliminating different types of backdoors, we conducted a preliminary study, which demonstrates that backdoors with low visibility can be easily destroyed by lightweight noise and those with high visibility need to be destroyed by high noise but can be easily detected. Based on the study, we propose SampDetox, which strategically combines lightweight and high noise. SampDetox applies weak noise to eliminate low-visibility backdoors and compares the structural similarity between the recovered and original samples to localize high-visibility backdoors. Intensive noise is then applied to these localized areas, destroying the high-visibility backdoors while preserving global semantic information. As a result, detoxified samples can be used for inference even by poisoned models. Comprehensive experiments demonstrate the
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- Test-Time Attention Purification for Backdoored Large Vision Language ModelsZhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 等CVPR 2026 · 被引用 7 次
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor AttackYukun Chen, Boheng Li, Yu Yuan, Leyi Qi 等NeurIPS 2025 · 被引用 6 次
- Enhancing Privacy in Multimodal Federated Learning with Information TheoryTianzhe Xiao, Yichen Li, Yining Qi, Yi Liu 等NeurIPS 2025 · 被引用 2 次
- BadTV: Unveiling Backdoor Threats in Third-Party Task VectorsChia-Yi Hsu, Yu-Lin Tsai, Zhe Yu, Yan-Lun Chen 等CCS 2026 · 被引用 2 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
相关 Paper
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan 等NeurIPS 2023 · 被引用 66 次
- Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution ShiftShengwei An, Sheng-Yen Chou, Kaiyuan Zhang, Qiuling Xu 等AAAI 2024 · 被引用 48 次
- SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction ConsistencyJunfeng Guo, Yiming Li, Xun Chen, Hanqing Guo 等ICLR 2023 · 被引用 19 次
- PBP: Post-training Backdoor Purification for Malware ClassifiersDung Thuy Nguyen, Ngoc N. Tran, Taylor T. Johnson, Kevin LeachNDSS 2025
- Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial PurificationGaozheng Pei, Shaojie Lyu, Gong Chen, Ke Ma 等CVPR 2025
