VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
Xinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, Min Zhang
Abstract
Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and predominantly feature videos rooted in Western cultural contexts. In this paper, we present VideoVista-CulturalLingo, the first video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension. Our work differs from existing benchmarks in the following ways: 1) Cultural diversity, incorporating cultures from China, North America, and Europe; 2) Multilinguistics, with questions presented in Chinese and English-two of the most widely spoken languages; and 3) Broad domain, featuring videos sourced from hundreds of humancreated domains. VideoVista-CulturalLingo contains 1,389 videos and 3,134 QA pairs, and we have evaluated 24 recent open-source or proprietary video large models. From the experiment results, we observe that: 1) Existing models perform worse on Chinese-centric questions than Western-centric ones, particularly those related to Chinese history; 2) Current open-source models still exhibit limitations in temporal understanding, especially in the Event Localization task, achieving a maximum score of only 45.2%; 3) Mainstream models demonstrate strong performance in general scientific questions, while open-source models demonstrate weak performance in mathematics. 1 * Corresponding author. 1 Evaluation Codes and Data are available at https:// videovista-culturallingo.github.io/ Question: 视频中主角一开始前往的寺庙的历史最早可以 追溯到哪个朝代? (To which dynasty can the history of the first temple mentioned in the video be traced?) A.唐朝(Tang Dynasty) B. 宋朝(Song Dynasty) C.明朝(Ming Dynasty) D. 清朝(Qing Dynasty)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a207811-eaa6-4787-86e2-d17e64962d94Builds on5
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Movie101: A New Movie Understanding BenchmarkZihao Yue, Qi Zhang, Anwen Hu, Liang Zhang et al.ACL 2023 · 11 citations
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.CVPR 2024
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary ResolutionZuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu et al.ICLR 2025
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language ModelsJiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu et al.ICLR 2025
Related papers
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang et al.AAAI 2025 · 7 citations
- CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingGuo Chen, Yicheng Liu, Yifei Huang, Baoqi Pei et al.ICLR 2025
- Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsZicheng Zhang, Ziheng Jia, Haoning Wu, Chunyi Li et al.CVPR 2025
- CURVE: A Benchmark for Cultural and Multilingual Long Video ReasoningDarshan Singh, Arsha Nagrani, Kawshik Manikantan, Harman Singh et al.CVPR 2026
