RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
Silpa Vadakkeeveetil Sreelatha, Sauradip Nag, Muhammad Awais, Serge J. Belongie, Anjan Dutta
摘要
The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this work, we propose RespoDiff, a novel framework for responsible text-to-image generation that incorporates a dual-module transformation on the intermediate bottleneck representations of diffusion models. Our approach introduces two distinct learnable modules: one focused on capturing and enforcing responsible concepts, such as fairness and safety, and the other dedicated to maintaining semantic alignment with neutral prompts. To facilitate the dual learning process, we introduce a novel score-matching objective that enables effective coordination between the modules. Our method outperforms state-of-the-art methods in responsible generation by ensuring semantic alignment while optimizing both objectives without compromising image fidelity. Our approach improves responsible and semantically coherent generation by 20% across diverse, unseen prompts. Moreover, it integrates seamlessly into large-scale models like SDXL, enhancing fairness and safety .The project page is available at https://vssilpa.github.io/respodiff_project_page.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept ControlBasim Azam, Naveed AkhtarCVPR 2025
- Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe GenerationSayedmoslem Shokrolahi, Jae-Mo Kang, Il-Min KimICML 2026
- PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image GenerationZongsheng Cao, Yangfan He, Anran Liu, Jun Xie 等ACM MM 2025 · 被引用 3 次
- Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image GenerationHang Li, Chengzhi Shen, Philip Torr, Volker Tresp 等CVPR 2024
- Mitigating Social Biases in Text-to-Image Diffusion Models via Linguistic-Aligned Attention GuidanceYue Jiang, Yueming Lyu, Ziwen He, Bo Peng 等ACM MM 2024 · 被引用 4 次
