ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep Generation
Yizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei, Qing Guo
摘要
The commercial text-to-image deep generation models ( e.g . DALL · E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent jailbreaking methods generate adversarial prompts capable of bypassing safety filters and producing unsafe content, exposing vulnerabilities in influential commercial models. However, once these adversarial prompts are identified, the safety filter can be updated to prevent the generation of unsafe images. In this work, we propose an effective, simple, and difficult-to-detect jailbreaking solution: generating safe content initially with normal text prompts and then editing the generations to embed unsafe content. The intuition behind this idea is that the deep generation model cannot reject safe generation with normal text prompts, while the editing models focus on modifying the local regions of images and do not involve a safety strategy. However, implementing such a solution is non-trivial, and we need to overcome several challenges: how to automatically confirm the normal prompt to replace the unsafe prompts, and how to effectively perform editable replacement and naturally generate unsafe content. In this work, we propose the collaborative generation and editing for jailbreaking text-to-image deep generation (ColJailBreak), which comprises three key components: adaptive normal safe substitution, inpainting-driven injection of unsafe content, and contrastive language-image-guided collaborative optimization. We validate our method on three datasets and compare it to two baseline methods. Our method
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Perception-Guided Jailbreak Against Text-to-Image ModelsYihao Huang, Le Liang, Tianlin Li, Xiaojun Jia 等AAAI 2025 · 被引用 34 次
- JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion ModelsXiaolong Jin, Zixuan Weng, Hanxi Guo, Chenlong Yin 等ICCV 2025 · 被引用 13 次
- Transstratal Adversarial Attack: Compromising Multi-Layered Defenses in Text-to-Image ModelsChunlong Xie, Kangjie Chen, Shangwei Guo, Shudong Zhang 等NeurIPS 2025 · 被引用 1 次
- ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned ModelsHyun Jun Yook, Ga San Jhun, Jae Hyun Cho, Min Jeon 等ICCV 2025 · 被引用 1 次
- STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity AttackXutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng 等ICML 2026
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Multimodal Pragmatic Jailbreak on Text-to-image ModelsTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang 等ACL 2025
- AdvPainting: Clean-text Jailbreaking Against Inpainting ModelsBingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan XuACM MM 2025
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong 等S&P 2024 · 被引用 188 次
- Modifier Unlocked: Jailbreaking Text-to-Image Models Through PromptsShuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong BaiS&P 2025
- JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution OptimizationHaolun Zheng, Yu He, Tailun Chen, Shuo Shao 等CVPR 2026 · 被引用 3 次
