Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
Joykirat Singh, Akshay Uttama Nambi, Vibhav Vineet
摘要
Large Language Models (LLMs) have been applied to Math Word Problems (MWPs) with transformative impacts, revolutionizing how these complex problems are approached and solved in various domains including educational settings. However, the evaluation of these models often prioritizes final accuracy, overlooking the crucial aspect of reasoning capabilities. This work addresses this gap by focusing on the ability of LLMs to detect and correct reasoning mistakes. We introduce a novel dataset MWP-MISTAKE, incorporating MWPs with both correct and incorrect reasoning steps generated through rule-based methods and smaller language models. Our comprehensive benchmarking reveals significant insights into the strengths and weaknesses of state-of-the-art models, such as GPT-4o, GPT-4, GPT-3.5Turbo, and others. We highlight GPT-4o's superior performance in mistake detection and rectification and the persistent challenges faced by smaller models. Additionally, we identify issues related to data contamination and memorization, impacting the reliability of LLMs in real-world applications. Our findings emphasize the importance of rigorous evaluation of reasoning processes and propose future directions to enhance the generalization and robustness of LLMs in mathematical problem-solving. Math Word Problems (MWPs) convey mathematical concepts and calculations through written descriptions, typically involving narrative scenarios [28] . Solvers must extract relevant mathematical information from these narratives and apply appropriate principles to arrive at solutions. Studies [34, 15, 11] have demonstrated that LLMs are proficient at understanding the contextual subtleties of MWPs, translating textual descriptions into mathematical expressions, and delivering precise solutions. Central to this process is mathematical reasoning, which enables models to adeptly manage complex, multi-step problems, draw logical inferences, and provide accurate solutions. Despite achieving remarkable accuracy rates exceeding 90% on datasets like GSM-8K (Grade School Math dataset with linguistically diverse word problems) [9] , foundational LLMs such as Claude-3-Opus [2], Gemini Ultra [29], and OpenAI reveal a significant gap in our understanding of their capabilities in mathematical reasoning [11] . Current research predominantly focuses on evaluating the final accuracy of MWPs [23, 35] , neglecting the intricate reasoning processes necessary to derive solutions. We argue that the reasoning steps play a pivotal role, and Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PromptHive: Bringing Subject Matter Experts Back to the Forefront with Collaborative Prompt Engineering for Educational Content CreationMohi Reza, Ioannis Anastasopoulos, Shreya Bhandari, Zachary A. PardosCHI 2025 · 被引用 18 次
- The Achilles’ Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language AbilitiesZixuan Qin, Qingchen Yu, Kunlin Lyu, Zhaoxin Fan 等ICLR 2026 · 被引用 10 次
- Are Reasoning LLMs Robust to Interventions on their Chain-of-Thought?Alexander von Recum, Leander Girrbach, Zeynep AkataICLR 2026 · 被引用 8 次
- Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal ThinkingYilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie 等ACL 2025
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
相关 Paper
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai 等ICLR 2025 · 被引用 3 次
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 等AAAI 2026 · 被引用 10 次
- It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetJiwoo Kim, Youngbin Kim, Ilwoong Baek, JinYeong Bak 等EMNLP 2023 · 被引用 2 次
- DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial DocumentsYilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi 等ACL 2024 · 被引用 8 次
