CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models
Ji Guo, xiaolong qin, Cencen Liu, Jielei Wang, Jierun Chen, Wenbo Jiang
Abstract
Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their applications become increasingly widespread, recent studies have revealed that VLMs are vulnerable to backdoor attacks. Existing backdoor attacks on VLMs primarily rely on data poisoning by adding visual triggers and modifying text labels, where the induced image–text mismatch makes poisoned samples easy to detect. To address this limitation, we propose the Clean-Label Backdoor Attack on VLMs via Diffusion Models (CBV), which leverages diffusion models to generate natural poisoned examples via score matching. Specifically, CBV modifies the score during the reverse generation process of the diffusion model to guide the generation of poisoned samples that contain triggered image features. To further enhance the effectiveness of the attack, we incorporate the textual information of the triggered images as multimodal guidance during generation. Moreover, to enhance stealthiness, we introduce a GradCAM-guided Mask (GM) that restricts modifications to only the most semantically important regions, rather than the entire image. We evaluate our method on MSCOCO and VQA v2 with four representative VLMs, achieving over 80% ASR while preserving normal functionality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb68007a-6596-4911-86bb-7c9de3e6535eCited by top-tier papers2
- DIVER: Diving Deeper into Distilled Data via Expressive Semantic RecoveryQianxin Xia, Zhiyong Shu, Wenbo Jiang, Jiawei Du et al.ICML 2026
- Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning ModelsJiakai Li, KE QIN, Rongzheng Wang, Yizhuo Ma et al.ICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Backdooring Vision-Language Models with Out-Of-Distribution DataWeimin Lyu, Jiachen Yao, Saumya Gupta, Lu Pang et al.ICLR 2025
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion ModelsChangkun Wu, Chenghao Chen, Wu kun, Chong Fu et al.CVPR 2026
- SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMsShuhan Xu, Siyuan Liang, Hongling Zheng, Aishan Liu et al.AAAI 2026 · 5 citations
- On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial MislabelingStanley Wu, Ronik Bhaskar, Anna Yoo Jeong Ha, Shawn Shan et al.CCS 2025
- TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language ModelsZhifang Zhang, Qiqi Tao, JIAQI LYU, Na Zhao et al.ICML 2026 · 5 citations
