GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image Models
Abhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper, Venkatesh Babu Radhakrishnan
摘要
Text-to-image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical diversity, reinforce stereotypes, and misrepresent regions. Given their broad reach, it is critical to rigorously evaluate how these models portray the world. Existing diversity metrics either rely on curated datasets or focus on surfacelevel visual similarity, limiting interpretability. We introduce GeoDiv, a framework leveraging large language and vision-language models to assess geographical diversity along two complementary axes: the Socio-Economic Visual Index (SEVI), capturing economic and condition-related cues, and the Visual Diversity Index (VDI), measuring variation in primary entities and backgrounds. Applied to images generated by models such as Stable Diffusion and FLUX.1-dev across 10 entities and 16 countries, GeoDiv reveals a consistent lack of diversity and identifies finegrained attributes where models default to biased portrayals. Strikingly, depictions of countries like India, Nigeria, and Colombia are disproportionately impoverished and worn, reflecting underlying socio-economic biases. These results highlight the need for greater geographical nuance in generative models. GeoDiv provides the first systematic, interpretable framework for measuring such biases, marking a step toward fairer and more inclusive generative systems. Project page: https: //abhipsabasu.github.io/geodiv * Equal contribution 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg 等ICLR 2024 · 被引用 139 次
- Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsAbhipsa Basu, R. Venkatesh Babu, Danish PruthiICCV 2023 · 被引用 54 次
相关 Paper
- OASIS Uncovers: High-Quality T2I Models, Same Old StereotypesSepehr Dehdashtian, Gautam Sreekumar, Vishnu BoddetiICLR 2025
- HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO DebiasingRuyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang 等ICML 2026 · 被引用 1 次
- Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image ModelsLili Zhang, Xi Liao, Zaijia Yang, Baihang Gao 等CHI 2024 · 被引用 17 次
- ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image GenerationAkshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo 等ACL 2024 · 被引用 18 次
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsAniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar 等ICCV 2025
