EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
Issa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke, Takanori Nakagawa, Kosuke Nakago, David Ha
摘要
Large Language Models (LLMs) have made remarkable progress, surpassing human performance on several benchmarks in domains such as mathematics and coding. A key driver of this progress has been the development of benchmark datasets. In contrast, the financial domain poses higher entry barriers due to its demand for specialized expertise, and benchmarks remain relatively scarce compared to those in mathematics or coding. We introduce EDINET-Bench, an open-source Japanese financial benchmark designed to evaluate LLMs on challenging tasks such as accounting fraud detection, earnings forecasting, and industry classification. EDINET-Bench is constructed from ten years of annual reports filed by Japanese companies. These tasks require models to process entire annual reports and integrate information across multiple tables and textual sections, demanding expert-level reasoning that is challenging even for human professionals. Our experiments show that even state-of-the-art LLMs struggle in this domain, performing only marginally better than logistic regression in binary classification tasks such as fraud detection and earnings forecasting. Our results show that simply providing reports to LLMs in a straightforward setting is not enough. This highlights the need for benchmark frameworks that better reflect the environments in which financial professionals operate, with richer scaffolding such as realistic simulations and task-specific reasoning support to enable more effective problem solving. We make our dataset and code publicly available to support future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set OverfittingTakashi Ishida, Thanawat Lodkaew, Ikko YamaneICML 2026 · 被引用 4 次
- Towards Scalable Oversight via Partitioned Human SupervisionRen Yin, Takashi Ishida, Masashi SugiyamaICLR 2026
它引用的顶会 Paper10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak 等ICLR 2024 · 被引用 220 次
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 被引用 165 次
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma 等EMNLP 2022 · 被引用 57 次
相关 Paper
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language ModelsDong Shu, Haoyang Yuan, Yuchen Wang, Yanguang Liu 等ACL 2026 · 被引用 11 次
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingZichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang 等ICCV 2025 · 被引用 1 次
- FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report GenerationSong Jin, Shuqi Li, Shukun Zhang, Rui YanAAAI 2026 · 被引用 1 次
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah 等EMNLP 2022 · 被引用 63 次
