FinMathBench: A Formula-Driven Benchmark for Evaluating LLMs' Math Reasoning Capabilities in Finance
Yi He, Ping Wang, Shiqiang Xiong, Chao Chen, Haixiang Hu
摘要
Many existing financial math reasoning benchmarks suffer from data contamination and high manual construction costs. To address this, we propose a novel formula-driven approach to dynamically construct math reasoning benchmarks in finance. Our two-stage approach: (1) generates single-formula questions by LLMs using a "Mask-for-Solve" paradigm for ground truth answers, and (2) synthesizes multi-formula questions through hierarchical tree-based DAGs. Our approach ensures novelty (via LLMs' creativity) and controllability of difficulty (via DAG structure). Based on a self-constructed financial formula bank, we utilize the proposed method to build FinMathBench, the first formula-driven and fully LLM-generated benchmark aimed at assessing LLMs' math reasoning abilities in finance, containing 946 questions across 4 complexity levels. Evaluation results on 40 LLMs demonstrate significant accuracy drops in multi-formula questions, e.g., 72.9% (1-Formula) to 14.0% (4-Formula) for GPT-4o under Chain-of-Thought prompting. Three critical flaws of LLMs are also observed: poor direct calculation performance, bias toward frequently solved variables in formulas, and erroneous "correction" of valid but extreme financial values. These findings highlight gaps in current LLMs' domain-specific reasoning and underscore FinMathBench's value for advancing robust financial LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Emergent and Predictable Memorization in Large Language ModelsStella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf 等NeurIPS 2023 · 被引用 205 次
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong 等ICLR 2024 · 被引用 92 次
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain 等NeurIPS 2021 · 被引用 76 次
相关 Paper
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingZichen Tang, Haihong E, Ziyan Ma, Haoyang He 等ACL 2025 · 被引用 17 次
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial ReasoningZhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan 等ACL 2026 · 被引用 13 次
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsChengke Zou, Xingang Guo, Rui Yang, Junyu Zhang 等ICLR 2025
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingZichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang 等ICCV 2025 · 被引用 1 次
