BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Xin Guo, Rongjunchen Zhang, Guilong Lu, Xuntao Guo, Jia Shuai, Zhi Yang, Liwen Zhang
摘要
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual efficacy in real-world scenarios. To tackle this challenge, we present BizFinBench.v2, the first integrated offline and online benchmark built upon authentic user query-response data from both Chinese and U.S. equity markets. It comprises 28,860 questions across eight offline and two online tasks. Experimental results show that GPT-5 achieves a mere 61.5% accuracy, still failing to meet the practical business requirement (84.8%). Among the evaluated commercial models, DeepSeek-R1 exhibits superior investment efficacy. Error analysis grounded in real financial practice reveals persistent limitations in existing models. By overcoming the constraints of prior benchmarks, BizFinBench.v2 provides a substantiated foundation for advancing LLM deployment in the financial sector. Our data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingYangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng 等NeurIPS 2024 · 被引用 197 次
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingZichen Tang, Haihong E, Ziyan Ma, Haoyang He 等ACL 2025 · 被引用 17 次
- FinQA: A Dataset of Numerical Reasoning over Financial DataZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah 等EMNLP 2021 · 被引用 8 次
- Dynalogue: A Transformer-Based Dialogue System with Dynamic AttentionRongjunchen Zhang, Tingmin Wu, Xiao Chen, Sheng Wen 等WWW 2023 · 被引用 4 次
- VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial UnderstandingZhaowei Liu, Xin Guo, Haotian Xia, Lingfeng Zeng 等EMNLP 2025
相关 Paper
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke 等ICLR 2026 · 被引用 9 次
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based AgentHaohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji 等ACL 2025
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingZichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang 等ICCV 2025 · 被引用 1 次
- FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report GenerationSong Jin, Shuqi Li, Shukun Zhang, Rui YanAAAI 2026 · 被引用 1 次
