Poisoning and Backdooring Contrastive Learning
Nicholas Carlini, Andreas Terzis
Abstract
Multimodal contrastive learning methods like CLIP train on noisy and uncurated training datasets. This is cheaper than labeling datasets manually, and even improves out-of-distribution robustness. We show that this practice makes backdoor and poisoning attacks a significant threat. By poisoning just 0.01% of a dataset (e.g., just 300 images of the 3 million-example Conceptual Captions dataset), we can cause the model to misclassify test images by overlaying a small patch. Targeted poisoning attacks, whereby the model misclassifies a particular test input with an adversarially-desired label, are even easier requiring control of 0.0001% of the dataset (e.g., just three out of the 3 million images). Our attacks call into question whether training on noisy and uncurated Internet scrapes is desirable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers97
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka et al.S&P 2024 · 309 citations
- Backdoor Defense via Decoupling the Training ProcessKunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin et al.ICLR 2022 · 253 citations
- BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised LearningJinyuan Jia, Yupei Liu, Neil Zhenqiang GongS&P 2022 · 200 citations
- CyCLIP: Cyclic Contrastive Language-Image PretrainingShashank Goel, Hritik Bansal, Sumit Bhatia, Ryan A. Rossi et al.NeurIPS 2022 · 192 citations
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Divide and Contrast: Self-supervised Learning from Uncurated DataYonglong Tian, Olivier J. Hénaff, Aäron van den OordICCV 2021 · 112 citations
Related papers
- Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanICML 2024 · 21 citations
- Detecting Backdoor Samples in Contrastive Language Image PretrainingHanxun Huang, Sarah Monazam Erfani, Yige Li, Xingjun Ma et al.ICLR 2025
- Robust Contrastive Language-Image Pretraining against Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanNeurIPS 2023 · 51 citations
- CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive LearningHritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover et al.ICCV 2023 · 78 citations
- BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive LearningSiyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu et al.CVPR 2024
