MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
Huanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, YuXin Song, Wenhao Wu, Dacheng Tao
Abstract
Reasoning plays a crucial role in advancing Multimodal Large Language Models (MLLMs) toward Artificial General Intelligence. However, existing MLLM benchmarks often fall short in precisely and comprehensively evaluating long-chain reasoning abilities from three key aspects: (1) lack of difficulty and diversity, (2) susceptibility to guessability and memorization, (3) inadequate assessment of intermediate reasoning steps. To fill this gap, we introduce MMReason, a new benchmark designed to precisely and comprehensively evaluate MLLM long-chain reasoning capability with diverse, open-ended, challenging questions. First, we curate challenging questions requiring multi-step reasoning from various fields (i.e., 6 disciplines) and multiple difficulty levels (i.e., from pre-university to university, and from foundational to competition tiers). Second, these questions are reformulated into an open-ended format and filtered using a multi-model voting technique to eliminate shortcut cases related to guessing and memorization, ensuring robust reasoning evaluations. Third, we annotate the questions with detailed step-by-step solutions, and design a reference-based ternary scoring mechanism to reliably assess intermediate reasoning steps. With MMReason, we benchmark popular leading MLLMs and provide an in-depth analysis of their reasoning capabilities. We hope MMReason will serve as a valuable resource for advancing MLLM reasoning research. Code will be available at https://github.com/HJYao00/MMReason.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70689c49-4f29-46e6-905d-fb2559e5e26dCited by top-tier papers7
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-Wise Group Relative Policy OptimizationJingyi Zhang, Jiaxing Huang, Huanjin Yao, Shunyu Liu et al.ICCV 2025 · 17 citations
- MM-DeepResearch: A Simple and Effective Multimodal Agentic Search BaselineHuanjin Yao, Qixiang Yin, Min Yang, Ziwang Zhao et al.ICML 2026 · 14 citations
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsJiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu et al.AAAI 2026 · 3 citations
- Benchmarking PhD-Level Coding in 3D Geometric Computer VisionWenyi Li, Renkai Luo, Yue Yu, Huan-ang Gao et al.CVPR 2026 · 2 citations
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu et al.ACL 2026 · 1 citation
Builds on18
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
Related papers
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsPengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu et al.AAAI 2026
- MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image ReasoningJiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang et al.ICLR 2026 · 7 citations
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research ProblemsXin Gui, King Zhu, JinCheng Ren, Qianben Chen et al.ICLR 2026 · 1 citation
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu et al.CVPR 2026 · 7 citations
