MRBench: A Multi-Image Reasoning Benchmark with Adaptive Knowledge Retrieval
Wenxi Huang, Xiaojun Chen, Qin Zhang, Ting Wan, Ziqi Liu, Liangjie Zhang
摘要
Multi-image understanding is crucial in real-world applications such as social media analysis and news reporting. However, existing benchmarks fall short in evaluating models' ability to integrate external knowledge and perform cross-image reasoning. To address this gap, we introduce MRBench, a comprehensive benchmark designed to assess knowledge-based reasoning across 12 diverse domains, incorporating four types of image relations: visually similar, identical entities, attribute-associated, and independent images. Additionally, we propose Multimodal Adaptive Retrieval Reasoning (MARR), a novel framework that enables the analysis of relationships among multiple input images and adaptively determines when to terminate the retrieval process. Extensive evaluations of state-of-the-art multimodal large language models (MLLMs) show a notable gap between model and human performance. The best-performing model, Gemini 2.0, reaches 56.86% accuracy, still 20.24% below humans. Proprietary models generally surpass open-source ones, particularly on visually similar and same-entity tasks, underscoring current limits in multi-image reasoning and retrieval and positioning MRBench as a key diagnostic tool. Our benchmark is available for further research and development in this field. https://github.com/Bruce-XJChen/MRBench.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?zhi zhu, YaoQi Fan, Zhe Chen, Yue Cao 等CVPR 2026
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li 等ICLR 2025
- MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image ReasoningJiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang 等ICLR 2026 · 被引用 7 次
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu 等ACL 2026 · 被引用 1 次
- MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal ModelsMingrui Wu, Hang Liu, Jiayi Ji, Xiaoshuai Sun 等CVPR 2026 · 被引用 5 次
