Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh
Abstract
Text-to-image diffusion models have recently enabled the creation of visually compelling, detailed images from textual prompts. However, their ability to accurately represent various cultural nuances remains an open question. In our work, we introduce CULTDIFF benchmark, evaluating whether state-of-the-art diffusion models can generate culturally specific images spanning ten countries. We show that these models often fail to generate cultural artifacts in architecture, clothing, and food, especially for underrepresented country regions, by conducting a fine-grained analysis of different similarity aspects, revealing significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images. With the collected human evaluations, we develop a neural-based image-image similarity metric, namely, CULTDIFF-S, to predict human judgment on real and generated images with cultural artifacts. Our work highlights the need for more inclusive generative AI systems and equitable dataset representation over a wide range of cultures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f32eeb9-8fdb-4e45-914f-a9702ac9c62cCited by top-tier papers3
- Culture in Action: Evaluating Text-to-Image Models through Social ActivitiesSina Malakouti, Boqing Gong, Adriana KovashkaICLR 2026 · 9 citations
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to ModelsLeander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell et al.ICLR 2026 · 5 citations
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language ModelsEunsu Kim, Junyeong Park, Na Min An, Junseong Kim et al.CVPR 2026 · 3 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image ModelsLili Zhang, Xi Liao, Zaijia Yang, Baihang Gao et al.CHI 2024 · 17 citations
- Synthetic History: Evaluating Visual Representations of the Past in Diffusion ModelsMaria-Teresa De Rosa Palmini, Eva CetinicICLR 2026 · 1 citation
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsAniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar et al.ICCV 2025
- PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance PredictionEduard Gabriel Poesina, Adriana Valentina Costache, Adrian-Gabriel Chifu, Josiane Mothe et al.CVPR 2025
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang et al.EMNLP 2024 · 6 citations
