PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Zhiqiang Wang, Xiaojun Jia
Abstract
Understanding the vulnerabilities of Large Vision Language Models (LVLMs) to jailbreak attacks is essential for their responsible realworld deployment. Most previous work requires access to model gradients, or is based on human knowledge (prompt engineering) to complete jailbreak, and they hardly consider the interaction of images and text, resulting in inability to jailbreak in black box scenarios or poor performance. To overcome these limitations, we propose a Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for toxicity maximization, referred to as PBI-Attack. Our method begins by extracting malicious features from a harmful corpus using an alternative LVLM and embedding these features into a benign image as prior information. Subsequently, we enhance these features through bidirectional cross-modal interaction optimization, which iteratively optimizes the bimodal perturbations in an alternating manner through greedy search, aiming to maximize the toxicity of the generated response. The toxicity level is quantified using a well-trained evaluation model. Experiments demonstrate that PBI-Attack outperforms previous state-of-the-art jailbreak methods, achieving an average attack success rate of 92.5% across three white-box LVLMs and around 67.3% on three black-box LVLMs. Our code is available at https://github.com/ Rosy0912/PBI-Attack . Disclaimer: This paper contains potentially disturbing and offensive content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80655fb1-736a-4aaf-b970-f9115bc4e415Cited by top-tier papers9
- Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsTeng Ma, Xiaojun Jia, Ranjie Duan, Xinfeng Li et al.ICCV 2025 · 35 citations
- Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM AlignmentRuoxi Cheng, Haoxuan Ma, Weixin Wang, Ranjie Duan et al.ICLR 2026 · 23 citations
- EcoAlign: An Economically Rational Framework for Efficient LVLM AlignmentRuoxi Cheng, Haoxuan Ma, Teng Ma, Hongyi ZhangCVPR 2026 · 6 citations
- ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play AttacksZhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang et al.ICLR 2026 · 5 citations
- MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMsYilian Liu, Guoshun Nan, Jiuyang Lyu, Zhican Chen et al.ICLR 2026 · 3 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao et al.AAAI 2026 · 4 citations
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 271 citations
- Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyShiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen et al.ICCV 2025 · 8 citations
- JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual SteeringRenmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan et al.ACM MM 2025 · 5 citations
- White-box Multimodal Jailbreaks Against Large Vision-Language ModelsRuofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji et al.ACM MM 2024 · 22 citations
