Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
Dadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He, Haoran Li, Yuxin Li, Yumeng Wang, Yi R. Fung
Abstract
Large reasoning models (e.g., R1, o3) have demonstrated remarkable mathematical problem-solving abilities. However, the high reported accuracy of these advanced models on popular datasets and reliance on purely numerical evaluation often mask their true reasoning shortcomings. To address this, we propose leveraging the inherent rigor and methodological complexity of mathematical proofs as a diagnostic tool to expose these hidden failures. Specifically, we introduce the RFMDataset (Reveal Failure Modes), a collection of 200 diverse mathematical proof problems to thoroughly evaluate the performance of advanced models. Our in-depth analysis of their failures uncovers 10 fine-grained error types, which shows fundamental limitations in current large reasoning models: 1) Large reasoning models still have limited capability in generating entirely correct mathematical proofs, with some models solving less than 20% of problems and even making mistakes on fundamental ones; 2) models exhibit a diverse spectrum of reasoning failures, prominently demonstrating the lack of guarantees for the correctness and rigor intermediate reasoning steps; and 3) models show hallucination and incompleteness during the reasoning process. Our findings also reveal that directly prompting models to self-reflect on specific failure modes is insufficient to resolve the current logical dilemmas, necessitating domain knowledge and formal verification. * Equal contribution. Question: Given a triangle ABC with side lengths AB, BC, and AC, prove that AB + BC > AC. Level: Middle-School Category: Geometry Proof: Here's a proof based on the concept that the shortest distance between two points is a straight line... 1. Consider the three points A, B, and C. These three points form the vertices of the triangle ABC. Model Experiment Evaluation Dataset Construction Let : [0, + ∞) → ℝ be a continuous function. For any ∈ [0, 1], we have lim →+∞ ( + ) = 0. Prove that lim →+∞ () may not exist. Here is a math proof question: Question. A student give an answer to this question. Here is his proof: Proof by LLM. And his thinking process: Thinking Process. You are a teacher need to justify whether the student's proof is right. You should examine the proof step by step. Here are possible mistakes:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92d91b95-65b8-47e6-862a-06e6951f597bCited by top-tier papers7
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong et al.ACL 2026 · 19 citations
- Diversity-Enhanced Reasoning for Subjective QuestionsYumeng Wang, Zhiyuan Fan, Jiayu Liu, Jen-Tse Huang et al.ICLR 2026 · 13 citations
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan et al.ACL 2026 · 5 citations
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren et al.ICML 2026 · 5 citations
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time ExplorationDadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian et al.ICLR 2026 · 3 citations
Builds on7
- The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical ProofsJasper Dekoninck, Ivo Petrov, Kristian Minchev, Miroslav Marinov et al.ICLR 2026 · 28 citations
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong et al.ACL 2026 · 19 citations
- Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math CapabilityRuida Wang, Yuxin Li, Yi R. Fung, Tong ZhangEMNLP 2025
- One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMsYinghui Li, Jiayi Kuang, Haojing Huang, Zhikun Xu et al.ICML 2025
- miniCTX: Neural Theorem Proving with (Long-)ContextsJiewen Hu, Thomas Zhu, Sean WelleckICLR 2025
Related papers
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary LinesYumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu et al.ACL 2026 · 3 citations
- IsarStep: a Benchmark for High-level Mathematical ReasoningWenda Li, Lei Yu, Yuhuai Wu, Lawrence C. PaulsonICLR 2021 · 69 citations
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang et al.ICLR 2026 · 14 citations
- RMath: A Logic Reasoning-Focused Datasets Toward Mathematical Multistep Reasoning TasksZiyi Hu, Jun Liu, Zhongzhi Liu, Yuzhong Liu et al.AAAI 2025 · 4 citations
