TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language Models
Hao Yu, Ke Liang, Junxian Duan, Jun Wang, Siwei Wang, Chuan Ma, Xinwang Liu
摘要
Large Vision-Language Models (LVLMs) enhance the capabilities of Large Language Models by integrating visual inputs, thereby enabling advanced multimodal reasoning across diverse applications. However, these enhanced reasoning capabilities introduce new security risks, particularly to jailbreaking attacks that bypass built-in safety mechanisms to elicit harmful or unauthorized outputs. While recent efforts have explored adversarial and typographic prompts, most existing attacks suffer from three key limitations: reliance on auxiliary models, limited effectiveness in black-box scenarios, and inadequate exploitation of the LVLMs' intrinsic reasoning abilities. In this work, we propose TVChain, a novel black-box jailbreaking framework that explicitly intervenes in both the visual and textual reasoning processes of LVLMs. TVChain decomposes malicious prompts into a sequence of semantically meaningful sub-images that represent relevant objects and behaviors, thereby circumventing direct exposure of illicit content. In parallel, a carefully designed chain-of-thought (CoT) textual prompt is employed to steer the model's reasoning toward reconstructing the intended activity in a covert yet effective manner. We demonstrate that this compositional prompting strategy reduces the likelihood of triggering safety mechanisms while preserving attack efficacy. Extensive evaluations on eleven LVLMs (seven open-source and four commercial) across two benchmark datasets and three state-of-the-art defenses validate the effectiveness and robustness of TVChain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 被引用 271 次
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa 等CHI 2024 · 被引用 81 次
- VLP: Vision Language Planning for Autonomous DrivingChenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik 等CVPR 2024 · 被引用 53 次
- An Image Is Worth 1000 Lies: Transferability of Adversarial Images across Prompts on Vision-Language ModelsHaochen Luo, Jindong Gu, Fengyuan Liu, Philip TorrICLR 2024 · 被引用 52 次
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li 等ICCV 2025 · 被引用 37 次
相关 Paper
- VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language ModelsBingrui Sima, Linhua Cong, Wenxuan Wang, Kun HeEMNLP 2025 · 被引用 11 次
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma 等ICLR 2026 · 被引用 4 次
- DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image InputsWenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang 等ACL 2026
- Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic BlueprintsChenxi Li, Xianggan Liu, Dake Shen, Yaosong Du 等CVPR 2026
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong 等NeurIPS 2025 · 被引用 14 次
