VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
Qilin Liao, Anamika Lochab, Ruqi Zhang
Abstract
Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplored vulnerabilities. Existing multimodal redteaming methods focus on a single modality while ignoring cross-modal interactions, rely heavily on handcrafted templates, and expose only a narrow subset of vulnerabilities. To address these limitations, we introduce VERA-V, a variational inference framework that recasts multimodal jailbreak discovery as learning a joint posterior distribution over paired text-image prompts. This probabilistic view captures complex cross-modal interactions, enabling stealthy, coordinated adversarial inputs that bypass model guardrails. We train a lightweight attacker to approximate the posterior, allowing efficient sampling of diverse jailbreaks and providing distributional insights into vulnerabilities. VERA-V further integrates three complementary strategies: (i) typographybased text prompts that embed harmful cues, (ii) diffusion-based image synthesis that introduces adversarial signals, and (iii) structured distractors to fragment VLM attention. Experiments on HarmBench and HADES benchmarks show that VERA-V consistently outperforms state-ofthe-art baselines on both open-source and frontier VLMs, achieving up to 53.75% higher attack success rate (ASR) over the best baseline on GPT-4o. We include the code on the project page available here: https://github.com/kxwhiowo/VERA-V Warning: This paper contains unfiltered content generated by VLMs that may be offensive to readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b162b77d-eb16-4d5e-bf5b-c02218e72ff1Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
Related papers
- VERA: Variational Inference Framework for Jailbreaking Large Language ModelsAnamika Lochab, Lu Yan, Patrick Pynadath, Xiangyu Zhang et al.NeurIPS 2025 · 3 citations
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma et al.ICLR 2026 · 4 citations
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong et al.NeurIPS 2025 · 14 citations
- Jailbreak Large Vision-Language Models Through Multi-Modal LinkageYu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang et al.ACL 2025 · 51 citations
- TRUST-VLM: Thorough Red-Teaming for Uncovering Safety Threats in Vision-Language ModelsKangjie Chen, Muyang Li, Guanlin Li, Shudong Zhang et al.ICML 2025
