Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar, Shivalika Singh, Fabian Farestam, Angelika Romanou, Danylo Boiko, Dipika Khullar, Mike Zhang, Dominik Krzeminski, Jekaterina Novikova
摘要
blend image and text modalities. Our dataset pushes beyond simple captioning tasks, challenging models to reason about visual content in various topics, the way humans are evaluated in exams worldwide. Through a large-scale open science effort across 18 languages, we construct Kaleidoscope (see Figure 1 ), featuring a diverse selection of knowledge domains across 14 subjects. With 55% of the total 20,911 questions requiring image understanding for accurate resolution, our work aims to establish a comprehensive, and inclusive evaluation framework for multimodal language models. We evaluate a wide range of state-of-the-art models on Kaleidoscope, including Claude 3.5 Sonnet (Anthropic, 2024), GPT-4o (OpenAI et al., 2024), and Gemini-V (Google et al., 2024), as well as smaller open-weight VLMs, such as Aya-Vision model family (Cohere-For-AI-Team, 2025), Molmo (Deitke et al., 2024) Pangea (Yue et al., 2025), and Qwen2.5-VL model family (Qwen-Team, 2025). Our key contributions and findings are highlighted here: et al., 2024), closely mimicking conventional human testing methodologies. Our work is built around three core design principles that guide the selection, curation, processing, and addition of exams: Multimodality: Images are central to Kaleidoscope, as we aim to evaluate how VLMs integrate and reason about visual information to answer questions. We prioritize multimodal questions with diverse image types, complemented by a similar proportion of text-only questions for a complete assessment and comparison. Multilinguality: The benchmark contains questions in 18 languages, with a focus on underrepresented mid-and low-resource languages (e.g., Nepali, Lithuanian) alongside high-resource languages (e.g., English, Spanish) for a thorough evaluation across a broad range of languages. Diversity: Our goal is to collect exams covering as wide a range of topics as possible ranging from Mathematics and Sociology, to Medicine and Driving Licenses, ensuring comprehensive evaluation across various domains. The final collection includes exams from 14 different domains, collected from 18 countries and with varying educational levels (from high school to professional exams), allowing detailed clustering and comprehensive evaluation. Global Collaboration Our work entailed an extensive, open science process to manually collect data by working directly with native speakers of different languages (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani 等ACL 2025 · 被引用 144 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy 等ICML 2022 · 被引用 71 次
相关 Paper
- EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language ModelsRocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Dimitrov 等ACL 2024 · 被引用 13 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long 等ICLR 2026 · 被引用 8 次
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla 等ICLR 2026 · 被引用 9 次
- Understanding ME? Multimodal Evaluation for Fine-grained Visual CommonsenseZhecan Wang, Haoxuan You, Yicheng He, Wenhao Li 等EMNLP 2022 · 被引用 2 次
