No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models
Angéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang, Andreas Steiner, Xiaohua Zhai, Ibrahim M. Alabdulmohsin
摘要
We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the common filtering of training data to English image-text pairs disadvantages communities of lower socioeconomic status and negatively impacts cultural understanding. Notably, this performance gap is not captured by - and even at odds with - the currently popular evaluation metrics derived from the Western-centric ImageNet and COCO datasets. Second, pretraining with global, unfiltered data before fine-tuning on English content can improve cultural understanding without sacrificing performance on said popular benchmarks. Third, we introduce the task of geo-localization as a novel evaluation metric to assess cultural diversity in VLMs. Our work underscores the value of using diverse data to create more inclusive multimodal systems and lays the groundwork for developing VLMs that better represent global perspectives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Meta CLIP 2: A Worldwide Scaling RecipeYung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh 等NeurIPS 2025 · 被引用 72 次
- Benchmarking Vision Language Models for Cultural UnderstandingShravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 等EMNLP 2024 · 被引用 26 次
- Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal SystemsIbrahim Alabdulmohsin, Xiaohua ZhaiNeurIPS 2025 · 被引用 8 次
- See It from My Perspective: How Language Affects Cultural Bias in Image UnderstandingAmith Ananthram, Elias Stengel-Eskin, Mohit Bansal, Kathleen McKeownICLR 2025 · 被引用 3 次
- GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image ModelsAbhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang 等EMNLP 2024 · 被引用 6 次
- Multilingual Diversity Improves Vision-Language RepresentationsThao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma 等NeurIPS 2024 · 被引用 19 次
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 等AAAI 2025 · 被引用 7 次
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla 等ICLR 2026 · 被引用 9 次
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang 等ICML 2025
