Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
Maria-Teresa De Rosa Palmini, Eva Cetinic
摘要
As Text-to-Image (TTI) diffusion models become increasingly influential in content creation, growing attention is being directed toward their societal and cultural implications. While prior research has primarily examined demographic and cultural biases, the ability of these models to accurately represent historical contexts remains largely underexplored. To address this gap, we introduce a benchmark for evaluating how TTI models depict historical contexts. The benchmark combines HistVis, a dataset of 30,000 synthetic images generated by three state-of-the-art diffusion models from carefully designed prompts covering universal human activities across multiple historical periods, with a reproducible evaluation protocol. We evaluate generated imagery across three key aspects: (1) Implicit Stylistic Associations: examining default visual styles associated with specific eras; (2) Historical Consistency: identifying anachronisms such as modern artifacts in pre-modern contexts; and (3) Demographic Representation: comparing generated racial and gender distributions against LLM-estimated historically plausible demographics. Our findings reveal systematic inaccuracies in historically themed generated imagery, as TTI models frequently stereotype past eras by incorporating unstated stylistic cues, introduce anachronisms, and fail to reflect plausible demographic patterns. By providing a reproducible benchmark for historical representation in generated imagery, this work provides an initial step toward building more historically accurate TTI models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- DALL-EVAL: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation ModelsJaemin Cho, Abhay Zala, Mohit BansalICCV 2023 · 被引用 283 次
- Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsAbhipsa Basu, R. Venkatesh Babu, Danish PruthiICCV 2023 · 被引用 54 次
- Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image ModelsLili Zhang, Xi Liao, Zaijia Yang, Baihang Gao 等CHI 2024 · 被引用 17 次
相关 Paper
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive?Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad 等ACL 2025
- Leveraging Diffusion Perturbations for Measuring Fairness in Computer VisionNicholas Lui, Bryan Chia, William Berrios, Candace Ross 等AAAI 2024 · 被引用 3 次
- Better Little People Pictures: Generative Creation of Demographically Diverse AnthropographicsPriya Dhawka, Lauren Perera, Wesley WillettCHI 2024 · 被引用 8 次
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image ModelsEslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan 等ICCV 2023 · 被引用 115 次
- Are Diffusion Models Vision-And-Language Reasoners?Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Chris Pal 等NeurIPS 2023 · 被引用 21 次
