DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial Documents
Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan
Abstract
Recent LLMs have demonstrated remarkable performance in solving exam-like math word problems. However, the degree to which these numerical reasoning skills are effective in real-world scenarios, particularly in expert domains, is still largely unexplored. This paper introduces DOCMATH-EVAL, a comprehensive benchmark specifically designed to evaluate the numerical reasoning capabilities of LLMs in the context of understanding and analyzing specialized documents containing both text and tables. We conduct an extensive evaluation of 48 LLMs using Chainof-Thought and Program-of-Thought prompting techniques, aiming to comprehensively assess the capabilities and limitations of existing LLMs in DOCMATH-EVAL. We found that even the current best-performing system (i.e., GPT-4o) still significantly lags behind human experts in solving complex numerical reasoning problems grounded in long contexts. We believe that DOCMATH-EVAL can serve as a valuable benchmark for evaluating LLMs' capabilities in solving challenging numerical reasoning problems within expert domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 178cfa6a-3267-40b5-8417-54d52b0e8a25Cited by top-tier papers9
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial ReasoningZhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan et al.ACL 2026 · 13 citations
- gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token ThrottlingTianyu Guo, Xianwei Zhang, Jiangsu Du, Zhiguang Chen et al.SC 2025 · 3 citations
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative RetrievalSubhendu Khatuya, Shashwat Naidu, Pawan Goyal, Niloy GangulyEMNLP 2025 · 3 citations
- FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang et al.EMNLP 2024 · 3 citations
- Table-R1: Inference-Time Scaling for Table Reasoning TasksZheyuan Yang, Lyuhao Chen, Arman Cohan, Yilun ZhaoEMNLP 2025 · 2 citations
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 168 citations
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu et al.ICLR 2023 · 41 citations
- TheoremQA: A Theorem-driven Question Answering DatasetWenhu Chen, Ming Yin, Max Ku, Pan Lu et al.EMNLP 2023 · 30 citations
Related papers
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
- Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with ChecklistZihao Zhou, Shudong Liu, Maizhen Ning, Wei Liu et al.ICLR 2025
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li et al.AAAI 2026 · 10 citations
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 36 citations
- CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive PerspectiveJiayu Liu, Zhenya Huang, Wei Dai, Cheng Cheng et al.ICML 2025
