Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows
Xiang Yang, Feifei Li, Mi Zhang, Geng Hong, Xiaoyu You, Mi Wen, Min Yang
Abstract
Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful content remains a critical challenge, particularly in image-to-image (I2I) editing tasks. Existing safety mechanisms are primarily designed for text-to-image (T2I) synthesis or U-Net-based architectures, which limits their effectiveness for unified safety mitigation in DiT-based frameworks. To bridge this gap, we propose Unified Visual Safety Regulator (UVR), a training-free safe generation framework that regulates unsafe semantics in generated images. UVR is grounded in an analysis of attention dynamics from the perspective of information flow in MM-Attn. We identify a task-independent start-up stage, during which unsafe semantics in output patches rapidly emerge and can be accurately localized, followed by task-specific semantic amplification and interference stages, where harmful signals are further propagated and entangled with benign content. Based on these observations, UVR mitigates unsafe generation through unified, targeted attention modulation and explicit restriction of harmful information flow over the identified unsafe output patches. Experiments across various concepts show that UVR achieves state-of-the-art safety performance by achieving 91% and 77% erase rate in image synthesis and editing tasks, while preserving visual quality and fidelity with minimal degradation. Code is available at https://github.com/deng12yx/UVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e78debb-cfa7-4451-bca5-0bc5bc4e57c2Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion TransformersChenyu Zhang, Lanjun Wang, Yueyang Cheng, Ruidong Chen et al.CCS 2026 · 1 citation
- SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow TransformersXiang Yang, Feifei Li, Mi Zhang, Geng Hong et al.CVPR 2026 · 2 citations
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image EditingJoonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim et al.ICCV 2025 · 5 citations
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin et al.ICCV 2025 · 2 citations
- T-Edit: Triple-Branch Diffusion Anchoring for Consistent EditingLinsong Shan, Laurence Yang, Zecan Yang, Shijie Lian et al.ICML 2026
