Modifier Unlocked: Jailbreaking Text-to-Image Models Through Prompts
Shuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong Bai
摘要
The unprecedented image generation capability of text-to-image models makes them double-edged swords. While these models allow users to create exquisite images through simple prompts, they also provide adversaries with opportunities to generate Not-Safe-for-Work (NSFW) content, referred to as the jailbreak attack. Despite built-in safety filters serving as a mitigation, their vulnerabilities and associated safety issues remain a significant concern. In this work, we propose MODX, the first modifier-based attack framework for jailbreaking text-to-image models. Modx leverages a heuristic algorithm with two heuristic functions (constraints) to identify modifiers that adjust the artistic genre to subtly introduce unsafe elements that drive the generated images towards NSFW. This approach takes advantage of the fact that filters are unlikely to reject images in certain styles or artistic forms, effectively inducing the models to generate NSFW content. We demonstrate the feasibility of modifier-based jailbreaking with a theoretical analysis, and provide experimental evidence of the effectiveness of MODX. Our results show that MODX outperforms existing methods in successfully achieving jailbreaking across four state-of-the-art text-to-image models. Moreover, we evaluate MODX across additional NSFW categories and on more models or model versions, demonstrating its strong scalability and generalization. Disclaimer: This paper contains NSFW language and imagery that could be offensive, distressing, and/or upsetting. Reader discretion is advised.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual SteganographySongze Li, Jiameng Cheng, Yiming Li, Xiaojun Jia 等NDSS 2026 · 被引用 9 次
- FracFace: Breaking the Visual Clues - Fractal-Based Privacy-Preserving Face RecognitionWanying Dai, Beibei Li, Naipeng Dong, Guangdong Bai 等NeurIPS 2025 · 被引用 7 次
- SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow TransformersXiang Yang, Feifei Li, Mi Zhang, Geng Hong 等CVPR 2026 · 被引用 2 次
- ReTrace: Reinforcement Learning-Guided Reconstruction Attacks on Machine UnlearningMengyao Ma, Shuofeng Liu, Minhui Xue, Surya Nepal 等ICLR 2026
- Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information FlowsXiang Yang, Feifei Li, Mi Zhang, Geng Hong 等ICML 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong 等S&P 2024 · 被引用 188 次
- Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsSong Yan, Hui Wei, Jinlong Fei, Guoliang Yang 等ACM MM 2025
- ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationYizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei 等NeurIPS 2024 · 被引用 22 次
- AdvPainting: Clean-text Jailbreaking Against Inpainting ModelsBingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan XuACM MM 2025
- Multimodal Pragmatic Jailbreak on Text-to-image ModelsTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang 等ACL 2025
