Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
Shuyang Hao, Bryan Hooi, Jun Liu, Kai-Wei Chang, Zi Huang, Yujun Cai
Abstract
Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis, we uncover two critical findings: scenariomatched images can significantly amplify harmful outputs, and contrary to common assumptions in gradient-based attacks, minimal loss values do not guarantee optimal attack effectiveness. Building on these insights, we introduce MLAI (Multi-Loss Adversarial Images), a novel jailbreak framework that leverages scenario-aware image generation for semantic alignment, exploits flat minima theory for robust adversarial image selection, and employs multi-image collaborative attacks for enhanced effectiveness. Extensive experiments demonstrate MLAI's significant impact, achieving attack success rates of 77.75% on MiniGPT-4 and 82.80% on LLaVA-2, substantially outperforming existing methods by margins of 34.37% and 12.77% respectively. Furthermore, MLAI shows considerable transferability to commercial black-box VLMs, achieving up to 60.11% success rate. Our work reveals fundamental visual vulnerabilities in current VLMs safety mechanisms and underscores the need for stronger defenses. Warning: This paper contains potentially harmful example text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f4acf1e-6522-4591-850e-52987228c9e2Cited by top-tier papers12
- Attention! Your Vision Language Model Could Be Maliciously ManipulatedXiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo et al.NeurIPS 2025 · 13 citations
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEctionRunqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li et al.CVPR 2026 · 13 citations
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma et al.ICLR 2026 · 4 citations
- Robustness of Vision Language Models Against Split-Image Harmful Input AttacksMd Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta MehnazCCS 2026 · 1 citation
- PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity MaximizationRuoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan et al.EMNLP 2025 · 1 citation
Builds on11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
Related papers
- White-box Multimodal Jailbreaks Against Large Vision-Language ModelsRuofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji et al.ACM MM 2024 · 22 citations
- Jailbreaking Vision-Language Models via Dissonance-Guided Suffix Optimization and Image-Phrase InjectionJiacheng Pi, Zhiguo Yang, Xingxing Huang, Dongsheng Xu et al.CVPR 2026
- Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsRylan Schaeffer, Dan Valentine, Luke Bailey, James Chua et al.ICLR 2025
- Ideator: Jailbreaking and Benchmarking Large Vision-Language Models Using ThemselvesRuofan Wang, Juncheng Li, Yixu Wang, Bo Wang et al.ICCV 2025 · 2 citations
- Jailbreaking Vision-Language Models Through the Visual ModalityAharon Azulay, Jan Dubiński, Zhuoyun Li, Atharv Mittal et al.ICML 2026 · 3 citations
