The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
Yixin Wan, Di Wu, Haoran Wang, Kai-Wei Chang
Abstract
Prompt-based "diversity interventions" are commonly adopted to improve the diversity of Text-to-Image (T2I) models depicting individuals with various racial or gender traits. However, will this strategy result in nonfactual demographic distribution, especially when generating real historical figures. In this work, we propose DemOgraphic FActualIty Representation (DoFaiR), a benchmark to systematically quantify the trade-off between using diversity interventions and preserving demographic factuality in T2I models. DoFaiR consists of 756 meticulously fact-checked test instances to reveal the factuality tax of various diversity prompts through an automated evidencesupported evaluation pipeline. Experiments on DoFaiR unveil that diversity-oriented instructions increase the number of different gender and racial groups in DALLE-3's generations at the cost of historically inaccurate demographic distributions. To resolve this issue, we propose Fact-Augmented Intervention (FAI), which instructs a Large Language Model (LLM) to reflect on verbalized or retrieved factual information about gender and racial compositions of generation subjects in history, and incorporate it into the generation context of T2I models. By orienting model generations using the reflected historical truths, FAI significantly improves the demographic factuality under diversity interventions while preserving diversity. OpenAI. 2023. Dall•e 3 system card.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 939f3571-2558-4a0e-8fd7-27880d227fdeCited by top-tier papers3
- Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMsAngelina Wang, Michelle Phan, Daniel E. Ho, Sanmi KoyejoACL 2025 · 17 citations
- MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation ModelsChejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie et al.ICLR 2025
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsAniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar et al.ICCV 2025
Builds on5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image ModelsEslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan et al.ICCV 2023 · 115 citations
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu et al.ICCV 2023 · 89 citations
- Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsWenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu et al.ICLR 2023 · 86 citations
Related papers
- F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality ConsiderationsTian Lan, Jiang Li, Yemin Wang, Xu Liu et al.EMNLP 2025 · 3 citations
- Synthetic History: Evaluating Visual Representations of the Past in Diffusion ModelsMaria-Teresa De Rosa Palmini, Eva CetinicICLR 2026 · 1 citation
- FairRAG: Fair Human Generation via Fair Retrieval AugmentationRobik Shrestha, Yang Zou, Qiuyu Chen, Zhiheng Li et al.CVPR 2024 · 6 citations
- HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO DebiasingRuyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang et al.ICML 2026 · 1 citation
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng et al.ICML 2024 · 3 citations
