The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
Yixin Wan, Di Wu, Haoran Wang, Kai-Wei Chang
摘要
Prompt-based "diversity interventions" are commonly adopted to improve the diversity of Text-to-Image (T2I) models depicting individuals with various racial or gender traits. However, will this strategy result in nonfactual demographic distribution, especially when generating real historical figures. In this work, we propose DemOgraphic FActualIty Representation (DoFaiR), a benchmark to systematically quantify the trade-off between using diversity interventions and preserving demographic factuality in T2I models. DoFaiR consists of 756 meticulously fact-checked test instances to reveal the factuality tax of various diversity prompts through an automated evidencesupported evaluation pipeline. Experiments on DoFaiR unveil that diversity-oriented instructions increase the number of different gender and racial groups in DALLE-3's generations at the cost of historically inaccurate demographic distributions. To resolve this issue, we propose Fact-Augmented Intervention (FAI), which instructs a Large Language Model (LLM) to reflect on verbalized or retrieved factual information about gender and racial compositions of generation subjects in history, and incorporate it into the generation context of T2I models. By orienting model generations using the reflected historical truths, FAI significantly improves the demographic factuality under diversity interventions while preserving diversity. OpenAI. 2023. Dall•e 3 system card.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMsAngelina Wang, Michelle Phan, Daniel E. Ho, Sanmi KoyejoACL 2025 · 被引用 17 次
- MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation ModelsChejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie 等ICLR 2025
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsAniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar 等ICCV 2025
它引用的顶会 Paper5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image ModelsEslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan 等ICCV 2023 · 被引用 115 次
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu 等ICCV 2023 · 被引用 89 次
- Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsWenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu 等ICLR 2023 · 被引用 86 次
相关 Paper
- F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality ConsiderationsTian Lan, Jiang Li, Yemin Wang, Xu Liu 等EMNLP 2025 · 被引用 3 次
- Synthetic History: Evaluating Visual Representations of the Past in Diffusion ModelsMaria-Teresa De Rosa Palmini, Eva CetinicICLR 2026 · 被引用 1 次
- FairRAG: Fair Human Generation via Fair Retrieval AugmentationRobik Shrestha, Yang Zou, Qiuyu Chen, Zhiheng Li 等CVPR 2024 · 被引用 6 次
- HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO DebiasingRuyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang 等ICML 2026 · 被引用 1 次
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng 等ICML 2024 · 被引用 3 次
