Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
Wenhan Yang, Jingdong Gao, Baharan Mirzasoleiman
摘要
Contrastive Language-Image Pre-training (CLIP) on large image-caption datasets has achieved remarkable success in zero-shot classification and enabled transferability to new domains. However, CLIP is extremely more vulnerable to targeted data poisoning and backdoor attacks, compared to supervised learning. Perhaps surprisingly, poisoning 0.0001% of CLIP pre-training data is enough to make targeted data poisoning attacks successful. This is four orders of magnitude smaller than what is required to poison supervised models. Despite this vulnerability, existing methods are very limited in defending CLIP models during pre-training. In this work, we propose a strong defense, SAFECLIP, to safely pre-train CLIP against targeted data poisoning and backdoor attacks. SAFECLIP warms up the model by applying unimodal contrastive learning (CL) on image and text modalities separately. Then, it divides the data into safe and risky sets, by applying a Gaussian Mixture Model to the cosine similarity of image-caption pair representations. SAFECLIP pre-trains the model by applying the CLIP loss to the safe set and applying unimodal CL to image and text modalities of the risky set separately. By gradually increasing the size of the safe set during pre-training, SAFECLIP effectively breaks targeted data poisoning and backdoor attacks without harming the CLIP performance. Our extensive experiments on CC3M, Visual Genome and MSCOCO demonstrate that SAFECLIP significantly reduces the success rate of targeted data poisoning attacks from 93.75% to 0% and that of various backdoor attacks from up to 100% to 0%, without harming CLIP's performance 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Defending Multimodal Backdoored Models by Repulsive Visual Prompt TuningZhifang Zhang, Shuo He, Haobo Wang, Bingquan Shen 等NeurIPS 2025 · 被引用 18 次
- CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset SeparationBinyan Xu, Fan Yang, Xilin Dai, Di Tang 等ACM MM 2025 · 被引用 12 次
- TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language ModelsZhifang Zhang, Qiqi Tao, JIAQI LYU, Na Zhao 等ICML 2026 · 被引用 5 次
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-trainingXin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo 等NeurIPS 2025 · 被引用 5 次
- InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation AnalysisMengyuan Sun, Yu Li, Yunjie Ge, Yuchen Liu 等NDSS 2026 · 被引用 3 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Robust Contrastive Language-Image Pretraining against Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanNeurIPS 2023 · 被引用 51 次
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and AlignmentTong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang 等EMNLP 2025 · 被引用 1 次
- Detecting Backdoor Samples in Contrastive Language Image PretrainingHanxun Huang, Sarah Monazam Erfani, Yige Li, Xingjun Ma 等ICLR 2025
- BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIPJiawang Bai, Kuofeng Gao, Shaobo Min, Shu-Tao Xia 等CVPR 2024
- CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive LearningHritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover 等ICCV 2023 · 被引用 78 次
