Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
Dadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He, Haoran Li, Yuxin Li, Yumeng Wang, Yi R. Fung
摘要
Large reasoning models (e.g., R1, o3) have demonstrated remarkable mathematical problem-solving abilities. However, the high reported accuracy of these advanced models on popular datasets and reliance on purely numerical evaluation often mask their true reasoning shortcomings. To address this, we propose leveraging the inherent rigor and methodological complexity of mathematical proofs as a diagnostic tool to expose these hidden failures. Specifically, we introduce the RFMDataset (Reveal Failure Modes), a collection of 200 diverse mathematical proof problems to thoroughly evaluate the performance of advanced models. Our in-depth analysis of their failures uncovers 10 fine-grained error types, which shows fundamental limitations in current large reasoning models: 1) Large reasoning models still have limited capability in generating entirely correct mathematical proofs, with some models solving less than 20% of problems and even making mistakes on fundamental ones; 2) models exhibit a diverse spectrum of reasoning failures, prominently demonstrating the lack of guarantees for the correctness and rigor intermediate reasoning steps; and 3) models show hallucination and incompleteness during the reasoning process. Our findings also reveal that directly prompting models to self-reflect on specific failure modes is insufficient to resolve the current logical dilemmas, necessitating domain knowledge and formal verification. * Equal contribution. Question: Given a triangle ABC with side lengths AB, BC, and AC, prove that AB + BC > AC. Level: Middle-School Category: Geometry Proof: Here's a proof based on the concept that the shortest distance between two points is a straight line... 1. Consider the three points A, B, and C. These three points form the vertices of the triangle ABC. Model Experiment Evaluation Dataset Construction Let : [0, + ∞) → ℝ be a continuous function. For any ∈ [0, 1], we have lim →+∞ ( + ) = 0. Prove that lim →+∞ () may not exist. Here is a math proof question: Question. A student give an answer to this question. Here is his proof: Proof by LLM. And his thinking process: Thinking Process. You are a teacher need to justify whether the student's proof is right. You should examine the proof step by step. Here are possible mistakes:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong 等ACL 2026 · 被引用 19 次
- Diversity-Enhanced Reasoning for Subjective QuestionsYumeng Wang, Zhiyuan Fan, Jiayu Liu, Jen-Tse Huang 等ICLR 2026 · 被引用 13 次
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan 等ACL 2026 · 被引用 5 次
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren 等ICML 2026 · 被引用 5 次
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time ExplorationDadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper7
- The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical ProofsJasper Dekoninck, Ivo Petrov, Kristian Minchev, Miroslav Marinov 等ICLR 2026 · 被引用 28 次
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong 等ACL 2026 · 被引用 19 次
- Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math CapabilityRuida Wang, Yuxin Li, Yi R. Fung, Tong ZhangEMNLP 2025
- One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMsYinghui Li, Jiayi Kuang, Haojing Huang, Zhikun Xu 等ICML 2025
- miniCTX: Neural Theorem Proving with (Long-)ContextsJiewen Hu, Thomas Zhu, Sean WelleckICLR 2025
相关 Paper
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary LinesYumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu 等ACL 2026 · 被引用 3 次
- IsarStep: a Benchmark for High-level Mathematical ReasoningWenda Li, Lei Yu, Yuhuai Wu, Lawrence C. PaulsonICLR 2021 · 被引用 69 次
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang 等ICLR 2026 · 被引用 14 次
- RMath: A Logic Reasoning-Focused Datasets Toward Mathematical Multistep Reasoning TasksZiyi Hu, Jun Liu, Zhongzhi Liu, Yuzhong Liu 等AAAI 2025 · 被引用 4 次
