Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability
Yingdong Shi, Changming Li, Yifan Wang, Yongxiang Zhao, Anqi Pang, Sibei Yang, Jingyi Yu, Kan Ren
摘要
Diffusion models have demonstrated impressive capabilities in synthesizing diverse content. However, despite their high-quality outputs, these models often perpetuate social biases, including those related to gender and race. These biases can potentially contribute to harmful realworld consequences, reinforcing stereotypes and exacerbating inequalities in various social contexts. While existing research on diffusion bias mitigation has predominantly focused on guiding content generation, it often neglects the intrinsic mechanisms within diffusion models that causally drive biased outputs. In this paper, we investigate the internal processes of diffusion models, identifying specific decision-making mechanisms, termed bias features, embedded within the model architecture. By directly manipulating these features, our method precisely isolates and adjusts the elements responsible for bias generation, permitting granular control over the bias levels in the generated content. Through experiments on both unconditional and conditional diffusion models across various social bias attributes, we demonstrate our method's efficacy in managing generation distribution while preserving image quality. We also dissect the discovered model mechanism, revealing different intrinsic features controlling fine-grained aspects of generation, boosting further research on mechanistic interpretability of diffusion models. The project website is at https://foundation-model-research.github.io/difflens .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma 等CVPR 2026 · 被引用 7 次
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned InterventionsAda Görgün, Fawaz Sammani, Nikos Deligiannis, Bernt Schiele 等ICLR 2026 · 被引用 7 次
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion ModelsKatarzyna Zaleska, Lukasz Popek, Monika Wysoczanska, Kamil DejaCVPR 2026 · 被引用 2 次
- BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image GenerationRajatsubhra Chakraborty, Xujun Che, Depeng Xu, Cori Faklaris 等KDD 2026 · 被引用 1 次
- The Latent Color Subspace: Emergent Order in High-Dimensional ChaosMateusz Pach, Jessica Bader, Quentin Bouniot, Serge Belongie 等ICML 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Fair Generation without Unfair Distortions: Debiasing Text-To-Image Generation with Entanglement-Free AttentionJeonghoon Park, Juyoung Lee, Chaeyeon Chung, Jaeseong Lee 等ICCV 2025 · 被引用 1 次
- FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent GuidanceMintong Kang, Vinayshekhar Bannihatti Kumar, Shamik Roy, Abhishek Kumar 等EMNLP 2025
- Exposing Hidden Biases in Text-to-Image Models via Automated Prompt SearchManos Plitsis, Giorgos Bouritsas, Vassilis Katsouros, Yannis PanagakisICML 2026
- Mitigating Social Biases in Text-to-Image Diffusion Models via Linguistic-Aligned Attention GuidanceYue Jiang, Yueming Lyu, Ziwen He, Bo Peng 等ACM MM 2024 · 被引用 4 次
- Fine-tuning Bias Neurons for Fair Text-to-Image GenerationFan Qi, Zhan Wang, Changsheng Xu, Huaiwen ZhangACM MM 2025
