RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Meng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang, Bolin Ni, Guo-Wei Yang
摘要
Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving, particularly in multi-disciplinary and multimodal contexts. In this paper, we introduce a graduate-level, multi-disciplinary, English-Chinese benchmark, dubbed as Reasoning Bench (R-Bench), for assessing the reasoning capability of both language and multimodal models. R-Bench spans 1,094 questions across 108 subjects for language model evaluation and 665 questions across 83 subjects for multimodal model testing in both English and Chinese. These questions are meticulously curated to ensure rigorous difficulty calibration, subject balance, and crosslinguistic alignment, enabling the assessment to be an Olympiad-level multi-disciplinary benchmark. We evaluate widely used models, including OpenAI o1, GPT-4o, DeepSeek-R1, etc. Experimental results indicate that advanced models perform poorly on complex reasoning, especially multimodal reasoning. Even the top-performing model OpenAI o1 achieves only 53.2% accuracy on our multimodal evaluation. Data and code are made publicly available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsYi Zhang, Bolin Ni, Xin-Sheng Chen, Hengrui Zhang 等ICLR 2026 · 被引用 30 次
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit MapsSicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong 等CVPR 2026 · 被引用 19 次
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering TaskSunqi Fan, Jiashuo Cui, Meng-Hao Guo, Shuojin YangNeurIPS 2025 · 被引用 15 次
- GenExam: A Multidisciplinary Text-to-Image ExamZhaokai Wang, Penghao Yin, Xiangyu Zhao, Changyao Tian 等ICML 2026 · 被引用 14 次
- Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data SynthesisXuanzhong Chen, Zile Qiao, Guoxin Chen, Liangcai Su 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsPengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu 等AAAI 2026
- R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal FormalizationYi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang 等ICCV 2025 · 被引用 21 次
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 等ACL 2024 · 被引用 18 次
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao 等ICLR 2025
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen 等ICCV 2025 · 被引用 4 次
