Jailbreaking Vision-Language Models via Dissonance-Guided Suffix Optimization and Image-Phrase Injection
Jiacheng Pi, Zhiguo Yang, Xingxing Huang, Dongsheng Xu, Ruizhi Zhong, Wenjie Ruan
Abstract
The integration of vision and language in Vision-Language Models (VLMs), while enabling multimodal capabilities, inherently expands their attack surface. Among existing white-box jailbreak methods, su!x-optimization-based approaches often rely on gradient approximations over discrete token spaces, yielding insu!cient guidance and causing optimization to stagnate in local optima, while imageperturbation-based ones frequently exhibit poor crossmodel transferability. In this work, we introduce DGSIP, a Dissonance-Guided Su!x Optimization and Image-Phrase Injection framework. DGSIP leverages predictive dissonance between the target model and an unaligned model to identify tokens suppressed by safety alignment, using them as a more e"ective signal than gradient-based cues for su!x optimization. It further reinforces the attack by jointly optimizing the content and presentation of phrase embedded in images to leverage VLMs' cross-modal sensitivity. Our extensive experiments demonstrate that DGSIP outperforms prior baselines across multiple safety benchmarks and a range of open-source VLMs (e.g., MiniGPT-4, InstructBlip and LLaVA). Notably, compared to baselines, our method exhibits much stronger transferability to commercial black-box VLMs, such as GPT-4o-Mini, Gemini 2.0 Flash and Qwen 2.5-VL. Based upon DGSIP, we empirically reveal critical vulnerabilities in the safeguard mechanisms of current VLMs, highlighting the need for more robust defense strategies. The implementation is available on https://github.com/Trusted-LLM/DGSIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 754fea1b-ee91-4c71-90ac-5aec4f83909cBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak AttacksYunhan Zhao, Xiang Zheng, Lin Luo, Yige Li et al.ICLR 2025
- Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsShuyang Hao, Bryan Hooi, Jun Liu, Kai-Wei Chang et al.CVPR 2025
- Robustness of Vision Language Models Against Split-Image Harmful Input AttacksMd Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta MehnazCCS 2026 · 1 citation
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyShiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen et al.ICCV 2025 · 8 citations
