Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
Shiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen, Caixin Kang, Shouwei Ruan, Jialing Tao, Yuefeng Chen, Hui Xue, Xingxing Wei
摘要
Multimodal Large Language Models (MLLMs) have achieved impressive performance and have been put into practical use in commercial applications, but they still have potential safety mechanism vulnerabilities. Jailbreak attacks are red teaming methods that aim to bypass safety mechanisms and discover MLLMs' potential risks. Existing MLLMs' jailbreak methods often bypass the model's safety mechanism through complex optimization methods or carefully designed image and text prompts. Despite achieving some progress, they have a low attack success rate on commercial closed-source MLLMs. Unlike previous research, we empirically find that there exists a Shuffle Inconsistency between MLLMs' comprehension ability and safety ability for the shuffled harmful instruction. That is, from the perspective of comprehension ability, MLLMs can understand the shuffled harmful text-image instructions well. However, they can be easily bypassed by the shuffled harmful instructions from the perspective of safety ability, leading to harmful responses. Then we innovatively propose a text-image jailbreak attack named SI-Attack. Specifically, to fully utilize the Shuffle Inconsistency and overcome the shuffle randomness, we apply a query-based black-box optimization method to select the most harmful shuffled inputs based on the feedback of the toxic judge model. A series of experiments show that SI-Attack can improve the attack's performance on three benchmarks. In particular, SI-Attack can obviously improve the attack success rate for commercial MLLMs such as GPT-4o or Claude-3.5-Sonnet. Warning: This paper contains examples of harmful texts and images, and reader discretion is recommended.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionZiqi Miao, Yi Ding, Lijun Li, Jing ShaoEMNLP 2025 · 被引用 21 次
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEctionRunqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li 等CVPR 2026 · 被引用 13 次
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying 等NeurIPS 2025 · 被引用 12 次
- Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective MemoryCe Zhang, Jinxi He, Junyi He, Katia Sycara 等CVPR 2026 · 被引用 5 次
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy ExplorationChunxiao Li, Lijun Li, Jing ShaoCVPR 2026 · 被引用 4 次
它引用的顶会 Paper14
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 等NeurIPS 2023 · 被引用 404 次
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang 等AAAI 2025 · 被引用 350 次
相关 Paper
- DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image InputsWenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang 等ACL 2026
- Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual SteganographySongze Li, Jiameng Cheng, Yiming Li, Xiaojun Jia 等NDSS 2026 · 被引用 9 次
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao 等AAAI 2026 · 被引用 4 次
- Reference Attack: A New Cross-Modal Jailbreaking Attack against Multimodal Large Language ModelsYulong Wang, Yifei Fu, Jiayi GaoACL 2026
- PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity MaximizationRuoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan 等EMNLP 2025 · 被引用 1 次
