Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang, Linyi Li, Hongxia Yang, Xiao Huang
Abstract
Recent studies have raised significant concerns regarding the reliability of current mathematical benchmarks, highlighting key limitations such as simplistic design and potential data contamination that undermine evaluation accuracy. Consequently, developing a reliable benchmark that effectively evaluates large language models' (LLMs) genuine capabilities in mathematical reasoning remains a critical challenge. To address these concerns, we propose RV-Bench, a novel evaluation methodology for Benchmarking LLMs with Random Variables in mathematical reasoning. Specifically, we develop question-generating functions to produce random variable questions (RVQs), whose background content mirrors the original benchmark problems, but with randomized variable combinations, rendering them "unseen" to LLMs. Models must completely understand the inherent question pattern to correctly answer RVQs with diverse variable combinations. Thus, an LLMs' genuine reasoning capability is reflected through its accuracy and robustness on RV-Bench. We conducted extensive experiments on over 30 representative LLMs across more than 1,000 RVQs. Our findings reveal that LLMs exhibit a proficiency imbalance between encountered and "unseen" data distributions. Furthermore, RV-Bench reveals that proficiency generalization across similar mathematical reasoning tasks is limited, but we verified that it can still be effectively elicited through test-time scaling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88824a90-d92c-4acd-bebb-3e5c99d4304eBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
Related papers
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
- BIG-Bench Extra HardMehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch et al.ACL 2025
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo et al.AAAI 2026 · 16 citations
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li et al.AAAI 2026 · 10 citations
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 163 citations
