Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
Yixin Wu, Yun Shen, Michael Backes, Yang Zhang
摘要
Text-to-image models, such as Stable Diffusion (SD), undergo iterative updates to improve image quality and address concerns such as safety. Improvements in image quality are straightforward to assess. However, how model updates resolve existing concerns and whether they raise new questions remain unexplored. This study takes an initial step in investigating the evolution of text-to-image models from the perspectives of safety, bias, and authenticity. Our findings, centered on Stable Diffusion, indicate that model updates paint a mixed picture. While updates progressively reduce the generation of unsafe images, the bias issue, particularly in gender, intensifies. We also find that negative stereotypes either persist within the same Non-White race group or shift towards other Non-White race groups through SD updates, yet with minimal association of these traits with the White race group. Additionally, our evaluation reveals a new concern stemming from SD updates: State-of-the-art fake image detectors, initially trained for earlier SD versions, struggle to identify fake images generated by updated versions. We show that fine-tuning these detectors on fake images generated by updated versions achieves at least 96.6% accuracy across various SD versions, addressing this issue. Our insights highlight the importance of continued efforts to mitigate biases and vulnerabilities in evolving text-to-image models. Disclaimer. This paper contains language and machinegenerated images that some readers may find offensive, disturbing, and/or distressing. Reader discretion is advised.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image ModelsAna-Maria Cretu, Klim Kireev, Amro Abdalla, Wisdom Obinna 等S&P 2026 · 被引用 4 次
- When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation ParadigmYe Leng, Junjie Chu, Mingjie Li, Chenhao Lin 等CVPR 2026 · 被引用 3 次
- Selective Fine-Tuning for Targeted and Robust Concept UnlearningMansi Mansi, Avinash Kori, Francesca Toni, Soteris DemetriouCCS 2026 · 被引用 2 次
- ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image ModelsDong Han, Yong LiICML 2026
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes 等USENIX Security 2025
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?Yunbo Lyu, Zhou Yang, Yuqing Niu, Jing Jiang 等ACM MM 2025 · 被引用 2 次
- FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable DiffusionGeorge Cazenavette, Avneesh Sud, Thomas Leung, Ben UsmanCVPR 2024
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Sha, Zheng Li, Ning Yu, Yang ZhangCCS 2023 · 被引用 123 次
- SustainDiffusion: Optimising the Social and Environmental Sustainability of Stable Diffusion ModelsGiordano d’Aloisio, Tosin Fadahunsi, Jay Choy, Rebecca Moussa 等ICSE 2026
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion ModelsByeonghu Na, Mina Kang, Jiseok Kwak, Minsang Park 等NeurIPS 2025 · 被引用 8 次
