Ideator: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang
Abstract
As large Vision-Language Models (VLMs) gain prominence, ensuring their safe deployment has become critical. Recent studies have explored VLM robustness against jailbreak at-tacks-techniques that exploit model vulnerabilities to elicit harmful outputs. However, the limited availability of diverse multimodal data has constrained current approaches to rely heavily on adversarial or manually crafted images derived from harmful text datasets, which often lack effectiveness and diversity across different contexts. In this paper, we propose IDEATOR, a novel jailbreak method that autonomously generates malicious image-text pairs for black-box jailbreak attacks. IDEATOR is grounded in the insight that VLMs themselves could serve as powerful red team models for generating multimodal jailbreak prompts. Specifically, IDEATOR leverages a VLM to create targeted jailbreak texts and pairs them with jailbreak images generated by a state-of-the-art diffusion model. Extensive experiments demonstrate IDEATOR's high effectiveness and transferability, achieving a 94% attack success rate (ASR) in jailbreaking MiniGPT-4 with an average of only 5.34 queries, and high ASRs of , and 75 % when transferred to LLaVA, InstructBLIP, and Chameleon, respectively. Building on IDEATOR's strong transferability and automated process, we introduce the VLJailbreakBench, a safety benchmark comprising 3,654 multimodal jailbreak samples. Our benchmark results on 11 recently released VLMs reveal significant gaps in safety alignment. For instance, our challenge set achieves ASRs of 46.31% on GPT40 and 19.65% on Claude-3.5-Sonnet, underscoring the urgent need for stronger defenses. Disclaimer: This paper contains content that may be disturbing or offensive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba350ae7-f88a-4daa-9b5c-b535bef7e236Cited by top-tier papers6
- Attention! Your Vision Language Model Could Be Maliciously ManipulatedXiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo et al.NeurIPS 2025 · 13 citations
- Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective MemoryCe Zhang, Jinxi He, Junyi He, Katia Sycara et al.CVPR 2026 · 5 citations
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma et al.ICLR 2026 · 4 citations
- Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMsJinqi Luo, Jinyu Yang, Tal Neiman, Lei Fan et al.CVPR 2026 · 1 citation
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyShiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen et al.ICCV 2025 · 8 citations
- Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsShuyang Hao, Bryan Hooi, Jun Liu, Kai-Wei Chang et al.CVPR 2025
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin et al.ACL 2025
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM SafetyJialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen et al.ICML 2026 · 5 citations
- White-box Multimodal Jailbreaks Against Large Vision-Language ModelsRuofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji et al.ACM MM 2024 · 22 citations
