BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning
Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, Ee-Chien Chang
Abstract
Studying backdoor attacks is valuable for model copyright protection and enhancing defenses. While existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario that backdoor attacks can remain effective even after defenses and introduces the BadCLIP attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder the backdoor unlearning through clean fine-tuning. Extensive experiments demonstrate that our attack significantly outperforms state-of-the-art baselines (+45.3% ASR) in the presence of SoTA backdoor defenses, rendering these mitigation and detection strategies virtually ineffective. Furthermore, our approach effectively attacks some more rigorous scenarios like downstream tasks. We believe that this paper raises awareness regarding the potential threats associated with the practical application of multimodal contrastive learning and encourages the development of more robust defense mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers40
- Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery DetectionJiawei Liang, Siyuan Liang, Aishan Liu, Xiaojun Jia et al.ICLR 2024 · 40 citations
- Breaking the False Sense of Security in Backdoor Defense through Re-Activation AttackMingli Zhu, Siyuan Liang, Baoyuan WuNeurIPS 2024 · 38 citations
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov et al.ICLR 2026 · 28 citations
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token GenerationRuoyu Chen, Xiaoqing Guo, Kangwei Liu, Siyuan Liang et al.CVPR 2026 · 19 citations
- Defending Multimodal Backdoored Models by Repulsive Visual Prompt TuningZhifang Zhang, Shuo He, Haobo Wang, Bingquan Shen et al.NeurIPS 2025 · 18 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li et al.ICCV 2021 · 639 citations
Related papers
- CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive LearningHritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover et al.ICCV 2023 · 78 citations
- BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIPJiawang Bai, Kuofeng Gao, Shaobo Min, Shu-Tao Xia et al.CVPR 2024
- Test-Time Multimodal Backdoor Detection by Contrastive PromptingYuwei Niu, Shuo He, Qi Wei, Zongyu Wu et al.ICML 2025
- InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation AnalysisMengyuan Sun, Yu Li, Yunjie Ge, Yuchen Liu et al.NDSS 2026 · 3 citations
- Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanICML 2024 · 21 citations
