GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong, Wei Bi
Abstract
Large language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks. However, there are increasing debates regarding whether these models truly understand and apply mathematical knowledge or merely rely on shortcuts for mathematical reasoning. One essential and frequently occurring evidence is that when the math questions are slightly changed, LLMs can behave incorrectly. This motivates us to evaluate the robustness of LLMs' math reasoning capability by testing a wide range of question variations. We introduce the adversarial grade school math (GSM-PLUS) dataset, an extension of GSM8K augmented with various mathematical perturbations. Our experiments on 25 LLMs and 4 prompting techniques show that while LLMs exhibit different levels of math reasoning abilities, their performances are far from robust. In particular, even for problems that have been solved in GSM8K, LLMs can make mistakes when new statements are added or the question targets are altered. We also explore whether more robust performance can be achieved by composing existing prompting methods, in which we try an iterative method that generates and verifies each intermediate thought based on its reasoning goal and calculation result.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e0fb27b-72c9-4be0-b004-7643c916a42bCited by top-tier papers54
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu et al.NeurIPS 2025 · 43 citations
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang et al.ICLR 2026 · 24 citations
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura et al.ICLR 2026 · 21 citations
- WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-trainingChangxin Tian, jiapeng wang, Qian Zhao, Kunlong Chen et al.ICLR 2026 · 20 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract ThinkingSilin Gao, Antoine Bosselut, Samy Bengio, Emmanuel AbbeICLR 2026 · 3 citations
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang et al.ICLR 2023 · 52 citations
- MathAttack: Attacking Large Language Models towards Math Solving AbilityZihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao et al.AAAI 2024 · 38 citations
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled BenchmarkMinglai Yang, Ethan Huang, Liang Zhang, Mihai Surdeanu et al.EMNLP 2025 · 2 citations
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsIman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel et al.ICLR 2025
