Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
Maria-Teresa De Rosa Palmini, Eva Cetinic
Abstract
As Text-to-Image (TTI) diffusion models become increasingly influential in content creation, growing attention is being directed toward their societal and cultural implications. While prior research has primarily examined demographic and cultural biases, the ability of these models to accurately represent historical contexts remains largely underexplored. To address this gap, we introduce a benchmark for evaluating how TTI models depict historical contexts. The benchmark combines HistVis, a dataset of 30,000 synthetic images generated by three state-of-the-art diffusion models from carefully designed prompts covering universal human activities across multiple historical periods, with a reproducible evaluation protocol. We evaluate generated imagery across three key aspects: (1) Implicit Stylistic Associations: examining default visual styles associated with specific eras; (2) Historical Consistency: identifying anachronisms such as modern artifacts in pre-modern contexts; and (3) Demographic Representation: comparing generated racial and gender distributions against LLM-estimated historically plausible demographics. Our findings reveal systematic inaccuracies in historically themed generated imagery, as TTI models frequently stereotype past eras by incorporating unstated stylistic cues, introduce anachronisms, and fail to reflect plausible demographic patterns. By providing a reproducible benchmark for historical representation in generated imagery, this work provides an initial step toward building more historically accurate TTI models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 650eaba0-8f04-439b-a6f8-a54dbce28eafBuilds on7
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- DALL-EVAL: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation ModelsJaemin Cho, Abhay Zala, Mohit BansalICCV 2023 · 283 citations
- Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsAbhipsa Basu, R. Venkatesh Babu, Danish PruthiICCV 2023 · 54 citations
- Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image ModelsLili Zhang, Xi Liao, Zaijia Yang, Baihang Gao et al.CHI 2024 · 17 citations
Related papers
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive?Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad et al.ACL 2025
- Leveraging Diffusion Perturbations for Measuring Fairness in Computer VisionNicholas Lui, Bryan Chia, William Berrios, Candace Ross et al.AAAI 2024 · 3 citations
- Better Little People Pictures: Generative Creation of Demographically Diverse AnthropographicsPriya Dhawka, Lauren Perera, Wesley WillettCHI 2024 · 8 citations
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image ModelsEslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan et al.ICCV 2023 · 115 citations
- Are Diffusion Models Vision-And-Language Reasoners?Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Chris Pal et al.NeurIPS 2023 · 21 citations
