JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion Models
Xiaolong Jin, Zixuan Weng, Hanxi Guo, Chenlong Yin, Siyuan Cheng, Guangyu Shen, Xiangyu Zhang
Abstract
Diffusion models are widely used in real-world applications, but ensuring their safety remains a major challenge. Despite many efforts to enhance the security of diffusion models, jailbreak and adversarial attacks can still bypass these defenses, generating harmful content. However, the lack of standardized evaluation makes it difficult to assess the robustness of diffusion model pipelines. To address this, we introduce JailbreakDiffBench, a comprehensive benchmark for systematically evaluating the safety of diffusion models against various attacks and under different defenses. Our benchmark includes a high-quality, humanannotated prompt and image dataset covering diverse attack scenarios. It consists of two key components: (1) an evaluation protocol to measure the effectiveness of moderation mechanisms and (2) an attack assessment module to benchmark adversarial jailbreak strategies. Through extensive experiments, we analyze existing filters and reveal critical weaknesses in current safety measures. JailbreakD-iffBench is designed to support both text-to-image and textto-video models, ensuring extensibility and reproducibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 087b9a3a-b415-4bcb-9aeb-b10306bab0b1Cited by top-tier papers2
- When Generative AI Is Intimate, Sexy, and Violent: Examining Not-Safe-For-Work (NSFW) Chatbots on FlowGPTXian Li, Yuanning Han, Di Liu, Pengcheng An et al.CHI 2026 · 1 citation
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li et al.ACL 2026
Builds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
Related papers
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language ModelsZifan Peng, Yule Liu, Zhen Sun, Mingchen Li et al.ICLR 2026 · 20 citations
- Multimodal Pragmatic Jailbreak on Text-to-image ModelsTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang et al.ACL 2025
- AdvPainting: Clean-text Jailbreaking Against Inpainting ModelsBingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan XuACM MM 2025
- JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution OptimizationHaolun Zheng, Yu He, Tailun Chen, Shuo Shao et al.CVPR 2026 · 3 citations
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language ModelsZirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li et al.ACL 2026 · 20 citations
