Benchmarking Vision Language Models for Cultural Understanding
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, Aishwarya Agrawal
摘要
Foundation models and vision-language pretraining have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed on general scene understandingrecognizing objects, attributes, and actions -rather than cultural comprehension. This study introduces CULTURALVQA, a visual question-answering benchmark aimed at assessing VLM's geo-diverse cultural understanding. We curate a collection of 2,378 image -question pairs with 1-5 answers per question representing cultures from 11 countries across 5 continents. The questions probe understanding of various facets of culture such as clothing, food, drinks, rituals, and traditions. Benchmarking VLMs on CULTURALVQA, including GPT-4o and Gemini, reveals disparity in their level of cultural understanding across regions, with strong cultural understanding capabilities for North America while significantly lower performance for Africa. We observe disparity in their performance across cultural facets too, with clothing, rituals, and traditions seeing higher performances than food and drink. These disparities help us identify areas where VLMs lack cultural understanding and demonstrate the potential of CULTURALVQA as a comprehensive evaluation set for gauging VLM progress in understanding diverse cultures. https://culturalvqa.org A p r , 2 0 2 3 Ja n , 2 0 2 4 A p r , 2 0 2 4 M a y , 2 0 2 4 Ju n , 2 0 2 4 20 30 40
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM CollaborationChaeHun Park, Yujin Baek, Jaeseok Kim, Yu-Jung Heo 等ACL 2025 · 被引用 17 次
- GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial TasksMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja 等ICCV 2025 · 被引用 11 次
- Culture in Action: Evaluating Text-to-Image Models through Social ActivitiesSina Malakouti, Boqing Gong, Adriana KovashkaICLR 2026 · 被引用 9 次
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture UnderstandingJiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi 等ICLR 2026 · 被引用 9 次
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park 等ACL 2025
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang 等EMNLP 2024 · 被引用 6 次
- Seeing Culture: A Benchmark for Visual Reasoning and GroundingBurak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried A. Mulyawan 等EMNLP 2025
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language ModelsEunsu Kim, Junyeong Park, Na Min An, Junseong Kim 等CVPR 2026 · 被引用 3 次
- No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsAngéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 等NeurIPS 2024 · 被引用 17 次
