Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried A. Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
摘要
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures. In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual options in the first stage are systematically organized into three types: those originating from the same country, those from different countries, or a mixed group. Notably, all options are derived from a singular category for each type. Progression to the second stage occurs only after a correct visual option is chosen. The SCB benchmark comprises 1,065 images that capture 138 cultural artifacts across five categories from seven Southeast Asia countries, whose diverse cultures are often overlooked, accompanied by 3,178 questions, of which 1,093 are unique and meticulously curated by human annotators. Our evaluation of various VLMs reveals the complexities involved in cross-modal cultural reasoning and highlights the disparity between visual reasoning and spatial grounding in culturally nuanced scenarios. The SCB serves as a crucial benchmark for identifying these shortcomings, thereby guiding future developments in the field of cultural reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng 等EMNLP 2021 · 被引用 32 次
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 被引用 31 次
相关 Paper
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang 等EMNLP 2024 · 被引用 6 次
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 等AAAI 2025 · 被引用 7 次
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu 等ACL 2026
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture UnderstandingJiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi 等ICLR 2026 · 被引用 9 次
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual ContextsTaebaek Hwang, Minseo Kim, Gisang Lee, Seonuk Kim 等EMNLP 2025
