MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Fanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu, Hao Tian, Tianshuo Yang, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao, Ping Luo, Kaipeng Zhang, Wenqi Shao
摘要
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development, moving us toward achieving sophisticated multimodal multi-image user interactions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park 等CVPR 2026 · 被引用 144 次
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language ReasoningYizhen Zhang, Yang Ding, Shuoshuo Zhang, Xinchen Zhang 等NeurIPS 2025 · 被引用 13 次
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual ChainsJuntian Zhang, Chuanqi Cheng, Yuhan Liu, Wei Liu 等ACL 2025 · 被引用 13 次
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang 等CVPR 2026 · 被引用 12 次
- MiCo: Multi-image Contrast for Reinforcement Visual ReasoningXi Chen, Mingkang Zhu, Shaoteng Liu, Xiaoyang Wu 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper27
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
相关 Paper
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li 等ICLR 2025
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
- MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi 等EMNLP 2024 · 被引用 7 次
- Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQAYue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan 等ACL 2024 · 被引用 1 次
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsPeng Xia, Siwei Han, Shi Qiu, Yiyang Zhou 等ICLR 2025
