Model-Contrastive Learning for Backdoor Elimination
Zhihao Yue, Jun Xia, Zhiwei Ling, Ming Hu, Ting Wang, Xian Wei, Mingsong Chen
Abstract
Due to the popularity of Artificial Intelligence (AI) techniques, we are witnessing an increasing number of backdoor injection attacks that are designed to maliciously threaten Deep Neural Networks (DNNs) causing misclassification. Although there exist various defense methods that can effectively erase backdoors from DNNs, they greatly suffer from both high Attack Success Rate (ASR) and a non-negligible loss in Benign Accuracy (BA). Inspired by the observation that a backdoored DNN tends to form a new cluster in its feature spaces for poisoned data, in this paper, we propose a novel two-stage backdoor defense method, named MCLDef, based on Model-Contrastive Learning (MCL). MCLDef can purify the backdoored model by pulling the feature representations of poisoned data towards those of their clean data counterparts. Due to the shrunken cluster of poisoned data, the backdoor formed by end-to-end supervised learning can be effectively eliminated. Comprehensive experimental results show that, with only 5% of clean data, MCLDef significantly outperforms state-of-the-art defense methods by up to 95.79% reduction in ASR, while in most cases, the BA degradation can be controlled within less than 2%. Our code is available at https://github.com/Zhihao151/MCL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5af98349-c8f5-4744-a1d4-db38dd0a692dCited by top-tier papers6
- FedMut: Generalized Federated Learning via Stochastic MutationMing Hu, Yue Cao, Anran Li, Zhiming Li et al.AAAI 2024 · 46 citations
- SampDetox: Black-box Backdoor Defense via Perturbation-based Sample DetoxificationYanxin Yang, Chentao Jia, Dengke Yan, Ming Hu et al.NeurIPS 2024 · 20 citations
- WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural NetworksJun Xia, Zhihao Yue, Yingbo Zhou, Zhiwei Ling et al.NeurIPS 2024 · 17 citations
- Prototype Guided Backdoor Defense via Activation Space ManipulationVenkat Adithya Amula, Sunayana Samavedam, Saurabh Saini, Avani Gupta et al.ICCV 2025 · 2 citations
- FilterFL: Knowledge Filtering-based Data-Free Backdoor Defense for Federated LearningYanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao et al.CCS 2025
Related papers
- Backdoor Defense via Decoupling the Training ProcessKunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin et al.ICLR 2022 · 253 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor AttacksJiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo et al.AAAI 2026
- Circumventing Backdoor Space via Weight SymmetryJie Peng, Hongwei Yang, Jing Zhao, Hengji Dong et al.ICML 2025
- Effective Backdoor Defense by Exploiting Sensitivity of Poisoned SamplesWeixin Chen, Baoyuan Wu, Haoqian WangNeurIPS 2022 · 129 citations
