MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, Jiaya Jia
Abstract
In this work, we introduce a novel evaluation 001 paradigm for Large Language Models, one that 002 challenges them to engage in meta-reasoning. 003 This approach addresses critical shortcomings 004 in existing math problem-solving benchmarks, 005 traditionally used to evaluate the cognitive capa-006 bilities of agents. Our paradigm shifts the focus 007 from result-oriented assessments, which often 008 overlook the reasoning process, to a more holis-009 tic evaluation that effectively differentiates the 010 cognitive capabilities among models. For ex-011 ample, in our benchmark, GPT-4 demonstrates 012 a performance five times better than GPT3.5. 013 The significance of this new paradigm lies in 014 its ability to reveal potential cognitive deficien-015 cies in LLMs that current benchmarks, such 016 as GSM8K, fail to uncover due to their satura-017 tion and lack of effective differentiation among 018 varying reasoning abilities. Our comprehen-019 sive analysis includes several state-of-the-art 020 math models from both open-source and closed-021 source communities, uncovering fundamental 022 deficiencies in their training and evaluation ap-023 proaches. 024 1 Introduction 025 Pretrained on trillions of tokens and possessed 026 with billions of parameters, today's large language 027 model (OpenAI, 2023; Anthropic, 2023; Touvron 028 et al., 2023) is capable of generating coherent texts 029 and achieved super-human performances in many 030 tasks (Bubeck et al., 2023; Hendrycks et al., 2021). 031 With the hope of differentiating different model's 032 cognitive ability, math questions are often selected 033 as a proxy evaluation task. However, despite the 034 complexity and diversity of these math problems, 035
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f08d7f2c-ba8c-4768-b3e9-e755cc3f408dCited by top-tier papers16
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsMingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou et al.ACL 2025 · 85 citations
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li et al.NeurIPS 2024 · 51 citations
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo et al.ICCV 2025 · 22 citations
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathShrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming et al.ACL 2026 · 13 citations
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process JudgesJiaxin Ai, Pengfei Zhou, Zhaopan Xu, Ming Li et al.ICCV 2025 · 9 citations
Builds on6
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang et al.EMNLP 2022 · 103 citations
- AlignBench: Benchmarking Chinese Alignment of Large Language ModelsXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang et al.ACL 2024 · 9 citations
Related papers
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang et al.NeurIPS 2025 · 11 citations
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
- CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive PerspectiveJiayu Liu, Zhenya Huang, Wei Dai, Cheng Cheng et al.ICML 2025
- Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with ChecklistZihao Zhou, Shudong Liu, Maizhen Ning, Wei Liu et al.ICLR 2025
- SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired BenchmarkYujie Hou, Mei Wang, Yaoyao Zhong, Ting Zhang et al.ACL 2026
