FinTextQA: A Dataset for Long-form Financial Question Answering
Jian Chen, Peilin Zhou, Yining Hua, Loh Xin, Kehui Chen, Ziyuan Li, Bing Zhu, Junwei Liang
Abstract
Accurate evaluation of financial questionanswering (QA) systems necessitates a comprehensive dataset encompassing diverse question types and contexts. However, current financial QA datasets lack scope diversity and question complexity. This work introduces FinTextQA, a novel dataset for long-form question answering (LFQA) in finance. FinTextQA comprises 1,262 high-quality, source-attributed QA pairs extracted and selected from finance textbooks and government agency websites.Moreover, we developed a Retrieval-Augmented Generation (RAG)-based LFQA system, comprising an embedder, retriever, reranker, and generator. A multi-faceted evaluation approach, including human ranking, automatic metrics, and GPT-4 scoring, was employed to benchmark the performance of different LFQA system configurations under heightened noisy conditions. The results indicate that: (1) Among all compared generators, Baichuan2-7B competes closely with GPT-3.5-turbo in accuracy score; (2) The most effective system configuration on our dataset involved setting the embedder, retriever, reranker, and generator as Ada2, Automated Merged Retrieval, Bge-Reranker-Base, and Baichuan2-7B, respectively; (3) models are less susceptible to noise after the length of contexts reaching a specific threshold. The dataset is publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9be181a8-850e-4255-995b-e4db3246325bCited by top-tier papers9
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial ReasoningZhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan et al.ACL 2026 · 13 citations
- AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor MiningHaochen Luo, Ho Tin Ko, Jiandong Chen, David Q. Sun et al.ICLR 2026 · 8 citations
- OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial DomainShuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong WenEMNLP 2025 · 6 citations
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand et al.ACL 2026 · 6 citations
- INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in InsuranceChenwei Lin, Hanjia Lyu, Xian Xu, Jiebo LuoICCV 2025 · 3 citations
Builds on12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang et al.ICLR 2024 · 468 citations
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang et al.EMNLP 2021 · 63 citations
Related papers
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial DomainSuifeng Zhao, Zhuoran Jin, Sujian Li, Jun GaoEMNLP 2025 · 1 citation
- RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question AnsweringRujun Han, Yuhao Zhang, Peng Qi, Yumo Xu et al.EMNLP 2024 · 10 citations
- NumCache: KV Cache Compression and Retrieval for Financial Document QAEftychia Makri, Peiwen Li, Yidong Jiang, Junrong Chen et al.KDD 2026
- Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in FinanceFengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang et al.ACM MM 2025 · 1 citation
- NitiBench: Benchmarking LLM Frameworks on Thai Legal Question Answering CapabilitiesPawitsapak Akarajaradwong, Pirat Pothavorn, Chompakorn Chaksangchaichot, Panuthep Tasawong et al.EMNLP 2025 · 1 citation
