DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial Documents
Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan
摘要
Recent LLMs have demonstrated remarkable performance in solving exam-like math word problems. However, the degree to which these numerical reasoning skills are effective in real-world scenarios, particularly in expert domains, is still largely unexplored. This paper introduces DOCMATH-EVAL, a comprehensive benchmark specifically designed to evaluate the numerical reasoning capabilities of LLMs in the context of understanding and analyzing specialized documents containing both text and tables. We conduct an extensive evaluation of 48 LLMs using Chainof-Thought and Program-of-Thought prompting techniques, aiming to comprehensively assess the capabilities and limitations of existing LLMs in DOCMATH-EVAL. We found that even the current best-performing system (i.e., GPT-4o) still significantly lags behind human experts in solving complex numerical reasoning problems grounded in long contexts. We believe that DOCMATH-EVAL can serve as a valuable benchmark for evaluating LLMs' capabilities in solving challenging numerical reasoning problems within expert domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial ReasoningZhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan 等ACL 2026 · 被引用 13 次
- gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token ThrottlingTianyu Guo, Xianwei Zhang, Jiangsu Du, Zhiguang Chen 等SC 2025 · 被引用 3 次
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative RetrievalSubhendu Khatuya, Shashwat Naidu, Pawan Goyal, Niloy GangulyEMNLP 2025 · 被引用 3 次
- FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang 等EMNLP 2024 · 被引用 3 次
- Table-R1: Inference-Time Scaling for Table Reasoning TasksZheyuan Yang, Lyuhao Chen, Arman Cohan, Yilun ZhaoEMNLP 2025 · 被引用 2 次
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 被引用 168 次
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu 等ICLR 2023 · 被引用 41 次
- TheoremQA: A Theorem-driven Question Answering DatasetWenhu Chen, Ming Yin, Max Ku, Pan Lu 等EMNLP 2023 · 被引用 30 次
相关 Paper
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
- Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with ChecklistZihao Zhou, Shudong Liu, Maizhen Ning, Wei Liu 等ICLR 2025
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 等AAAI 2026 · 被引用 10 次
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 被引用 36 次
- CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive PerspectiveJiayu Liu, Zhenya Huang, Wei Dai, Cheng Cheng 等ICML 2025
