VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
Yueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li, Graham Neubig, Xiang Yue
摘要
Current multimodal benchmarks often conflate reasoning with domainspecific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings. To address this, we introduce VISUALPUZZLES, a benchmark that targets visual reasoning while deliberately minimizing reliance on specialized knowledge. VISUALPUZZLES consists of diverse questions spanning five categories: algorithmic, analogical, deductive, inductive, and spatial reasoning. One major source of our questions is manually translated logical reasoning questions from the Chinese Civil Service Examination. Experiments show that VISUALPUZ-ZLES requires significantly less intensive domain-specific knowledge and more complex reasoning compared to benchmarks like MMMU, enabling us to better evaluate genuine multimodal reasoning. Evaluations show that state-of-the-art multimodal large language models consistently lag behind human performance on VISUALPUZZLES, and that strong performance on knowledge-intensive benchmarks does not necessarily translate to success on reasoning-focused, knowledge-light tasks. Additionally, reasoning enhancements such as scaling up inference compute (with "thinking" modes) yield inconsistent gains across models and task types, and we observe no clear correlation between model size and performance. We also found that models exhibit different reasoning and answering patterns on VISU-ALPUZZLES compared to benchmarks with heavier emphasis on knowledge. VISUALPUZZLES offers a clearer lens through which to evaluate reasoning capabilities beyond factual recall and domain knowledge. * Equal Contributions. † Equal Contributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit MapsSicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong 等CVPR 2026 · 被引用 19 次
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang 等CVPR 2026 · 被引用 16 次
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?Yang Chen, Minghao Liu, Yufan Shen, Yunwen Li 等ICLR 2026 · 被引用 11 次
- SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMsSiting Wang, Minnan Pei, Luoyang Sun, Cheng Deng 等ICLR 2026 · 被引用 8 次
- MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image ReasoningJiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper8
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 等ACL 2025 · 被引用 377 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- VASR: Visual Analogies of Situation RecognitionYonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf 等AAAI 2023 · 被引用 27 次
- Are Deep Neural Networks SMARTer Than Second Graders?Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith 等CVPR 2023
相关 Paper
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language ModelsZesen Lyu, Dandan Zhang, Wei Ye, Fangdi Li 等EMNLP 2025
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu 等ICML 2026
