On Evaluating Adversarial Robustness of Large Vision-Language Models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, Min Lin
Abstract
Large vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models such as ChatGPT. Nonetheless, multimodal generation exacerbates safety concerns, since adversaries may successfully evade the entire system by subtly manipulating the most vulnerable modality (e.g., vision). To this end, we propose evaluating the robustness of open-source large VLMs in the most realistic and high-risk setting, where adversaries have only black-box system access and seek to deceive the model into returning the targeted responses. In particular, we first craft targeted adversarial examples against pretrained models such as CLIP and BLIP, and then transfer these adversarial examples to other VLMs such as MiniGPT-4, LLaVA, UniDiffuser, BLIP-2, and Img2Prompt. In addition, we observe that black-box queries on these VLMs can further improve the effectiveness of targeted evasion, resulting in a surprisingly high success rate for generating targeted responses. Our findings provide a quantitative understanding regarding the adversarial vulnerability of large VLMs and call for a more thorough examination of their potential security flaws before deployment in practice. Code is at https://github.com/yunqing-me/AttackVLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dbf8ee1-8b7d-4c0d-b926-725388e74527Cited by top-tier papers144
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang et al.AAAI 2025 · 350 citations
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 271 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- Image Hijacks: Adversarial Images can Control Generative Models at RuntimeLuke Bailey, Euan Ong, Stuart Russell, Scott EmmonsICML 2024 · 171 citations
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du et al.ICML 2024 · 128 citations
Builds on38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language ModelsChristian Schlarmann, Naman Deep Singh, Francesco Croce, Matthias HeinICML 2024 · 114 citations
- Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial AttacksPeng Xie, Yequan Bie, Jianda Mao, Yangqiu Song et al.CVPR 2025
- Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language ModelsJiaming Zhang, Junhong Ye, Xingjun Ma, Yige Li et al.CVPR 2025
- Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video ApproachLinhao Huang, Xue Jiang, Zhiqiang Wang, Wentao Mo et al.AAAI 2026 · 6 citations
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen et al.CVPR 2025
