DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor Attacks
Jiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo, Kun Fang, Ruochen Du, Yingnan Zhao, Haibo Hu
Abstract
Contrastive learning (CL) is a popular learning paradigm that excels in extracting meaningful representations from unlabeled data. Recent studies have shown that CL is highly vulnerable to backdoor attacks. Current defenses against backdoor attacks in CL are primarily reactive and post-training. That is, the detection and elimination of backdoors are executed in the deployment phase of a given well-trained model. However, these post-training defenses are usually prone to degrading model utility and resource-intensive, causing that the backdoor detection and elimination from a fully-trained model is quite challenging. To address this issue, we argue for a fundamental perspective, i.e., integrating the defense into the model's training phase, and propose a novel framework to mitigate the backdoor in CL, namely Density-Based Identification and Fine-Tuning (DIFT). Specifically, DIFT identifies potential poisoned samples during the early training phase via detecting embeddings with abnormal poisoning characteristic in the feature space. Then, to remove backdoors and preserve model utility, the detected poisoned samples are leveraged to fine-tune the model, and the remaining clean samples are further involved into training the model after the fine-tuning. DIFT, as a proactive training-time defense, avoids the problematic backdoor removal and the high computational cost associated with those reactive post-training methods. We empirically evaluate DIFT on various CL algorithms against backdoor attack. Experimental results demonstrate that our method exhibits promising defense effectiveness while maintaining model's clean data accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ff41cb4-6183-4f5b-b6d1-846ef9022786Builds on22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Blind Backdoors in Deep Learning ModelsEugene Bagdasaryan, Vitaly ShmatikovUSENIX Security 2021 · 372 citations
- Adversarial Unlearning of Backdoors via Implicit HypergradientYi Zeng, Si Chen, Won Park, Zhuoqing Mao et al.ICLR 2022 · 235 citations
Related papers
- Model-Contrastive Learning for Backdoor EliminationZhihao Yue, Jun Xia, Zhiwei Ling, Ming Hu et al.ACM MM 2023 · 9 citations
- Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language ModelsPeihai Jiang, Xixiang Lyu, Yige Li, Jing MaAAAI 2025 · 8 citations
- Backdoor Mitigation by Distance-Driven DetoxificationShaokui Wei, Jiayin Liu, Hongyuan ZhaICCV 2025 · 1 citation
- Effective Backdoor Defense by Exploiting Sensitivity of Poisoned SamplesWeixin Chen, Baoyuan Wu, Haoqian WangNeurIPS 2022 · 129 citations
- Data Poisoning Based Backdoor Attacks to Contrastive LearningJinghuai Zhang, Hongbin Liu, Jinyuan Jia, Neil Zhenqiang GongCVPR 2024 · 12 citations
