Highly Transferable Diffusion-based Unrestricted Adversarial Attack on Pre-trained Vision-Language Models
Wenzhuo Xu, Kai Chen, Ziyi Gao, Zhipeng Wei, Jingjing Chen, Yu-Gang Jiang
Abstract
Pre-trained Vision-Language Models (VLMs) have shown great ability in various Vision-Language tasks. However, these VLMs exhibit inherent vulnerabilities to transferable adversarial examples, which could potentially undermine their performance and reliability in real-world applications. Cross-modal interactions have been demonstrated to be the key point to boosting adversarial transferability, but the utilization of them is limited in existing multimodal adversarial attacks. Stable Diffusion, which contains multiple cross-attention modules, possesses great potential in facilitating adversarial transferability by leveraging abundant cross-modal interactions. Therefore, We propose a Multimodal Diffusion-based Attack (MDA), which conducts adversarial attacks against VLMs using Stable Diffusion. Specifically, MDA initially generates adversarial text, which is subsequently utilized to optimize the adversarial image during the diffusion process. Besides leveraging adversarial text in calculating downstream loss, MDA also takes it as the guiding prompt in adversarial image generation during the denoising process, which enriches the ways of cross-modal interactions, thus strengthening the adversarial transferability. Compared with pixel-based attacks, MDA introduces perturbations in the latent space rather than pixel space to manipulate high-level semantics, which is also beneficial to improving adversarial transferability. Experimental results demonstrate that the adversarial examples generated by MDA are highly transferable across different VLMs on different downstream tasks, surpassing state-of-the-art methods by a large margin.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9893c819-bd92-4e9a-8d36-fe40d9dbcdb8Cited by top-tier papers9
- Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented GenerationYingjia Shang, Yi Liu, Huimin Wang, Furong Li et al.KDD 2026 · 2 citations
- GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial PerturbationsXinwei Liu, Xiaojun Jia, Yuan Xun, Simeng Qin et al.AAAI 2026 · 2 citations
- Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning ModelsJiaming Zhang, Che Wang, Yang Cao, Longtao Huang et al.ICLR 2026 · 1 citation
- Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text UpdatesJaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee KimACL 2025
- From Zero to Hero: Cross-modal-enhanced Adversarial Item Promotion Attack against Multimodal Recommender SystemsMengyu Yao, Ziqi Zhang, Yifeng Cai, Junlin Liu et al.USENIX Security 2026
Related papers
- Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training ModelsDong Lu, Zhiqiang Wang, Teng Wang, Weili Guan et al.ICCV 2023 · 141 citations
- Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive InteractionYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou et al.CVPR 2026
- Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial AttacksPeng Xie, Yequan Bie, Jianda Mao, Yangqiu Song et al.CVPR 2025
- GLEAM: Enhanced Transferable Adversarial Attacks for Vision-Language Pre-Training Models via Global-Local TransformationsYunqi Liu, Xue Ouyang, Xiaohui CuiICCV 2025 · 9 citations
- Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training ModelsYang Li, Jia-Li Yin, Luojun Lin, Wei LinCVPR 2026
