From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang, Vered Shwartz
Abstract
Despite recent advancements in visionlanguage models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts. To address these limitations, we introduce the GLOBALRG benchmark, comprising two challenging tasks: retrieval across universals and cultural visual grounding. The former task entails retrieving culturally diverse images for universal concepts from 50 countries, while the latter aims at grounding culture-specific concepts within images from 15 countries. Our evaluation across a wide range of models reveals that the performance varies significantly across cultures -underscoring the necessity for enhancing multicultural understanding in vision-language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4e3e999-789b-49bf-930a-ac8788d664d2Cited by top-tier papers8
- Benchmarking Vision Language Models for Cultural UnderstandingShravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy et al.EMNLP 2024 · 26 citations
- Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM CollaborationChaeHun Park, Yujin Baek, Jaeseok Kim, Yu-Jung Heo et al.ACL 2025 · 17 citations
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture UnderstandingJiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi et al.ICLR 2026 · 9 citations
- Advancing Social Intelligence in AI Agents: Technical Challenges and Open QuestionsLeena Mathur, Paul Pu Liang, Louis-Philippe MorencyEMNLP 2024 · 6 citations
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally AwarenessYuchen Song, Andong Chen, Wenxin Zhu, Kehai Chen et al.ICLR 2026 · 3 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
Related papers
- No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsAngéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang et al.NeurIPS 2024 · 17 citations
- Seeing Culture: A Benchmark for Visual Reasoning and GroundingBurak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried A. Mulyawan et al.EMNLP 2025
- See It from My Perspective: How Language Affects Cultural Bias in Image UnderstandingAmith Ananthram, Elias Stengel-Eskin, Mohit Bansal, Kathleen McKeownICLR 2025 · 3 citations
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng et al.EMNLP 2021 · 32 citations
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang et al.AAAI 2025 · 7 citations
