OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
Shuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong Wen
Abstract
Retrieval-augmented generation (RAG) has emerged as a key application of large language models (LLMs), especially in vertical domains where LLMs lack domain-specific knowledge. Nevertheless, current RAG benchmarks often suffer from narrow scenarios and limited evaluation dimensions, hindering an all-sides understanding of RAG models in real-world vertical applications. This paper introduces Om-niEval, an omnidirectional and automatic RAG benchmark for the financial domain, featured by its omnidirectional evaluation framework: First, we categorize RAG scenarios by five task classes and 16 financial topics, leading to a matrix-based structured assessment. Next, we leverage a multi-dimensional and auto-chained data generation pipeline that integrates LLMbased automatic generation and human annotation approaches, creating high-quality evaluation instances. Further, we adopt a multi-stage evaluation to assess both retrieval and generation performance, resulting in a holistic RAG evaluation. Finally, rule-based and LLM-based metrics are combined to build a multi-level evaluation system. Our experiments indicate that the performance of RAG systems varies across topics and tasks, highlighting the importance of multi-aspect and structured assessments to better locate the advantages and disadvantages of RAG systems. We release our code at https://github.com/RUC-NLPIR/OmniEval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 759d473c-789b-468c-9add-cfb38da04b4fCited by top-tier papers8
- HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web SearchesJiejun Tan, Zhicheng Dou, Yan Yu, Jiehan Cheng et al.AAAI 2026 · 5 citations
- RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation SystemsShaobo Li, Yirui Zhou, Yuan Xu, Kevin Chen et al.VLDB 2026 · 3 citations
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG BenchmarksLiyang He, Yuren Zhang, Ziwei Zhu, Zhenghui Li et al.ICLR 2026 · 1 citation
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial DomainSuifeng Zhao, Zhuoran Jin, Sujian Li, Jun GaoEMNLP 2025 · 1 citation
- Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in FinanceFengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang et al.ACM MM 2025 · 1 citation
Builds on5
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah et al.EMNLP 2022 · 63 citations
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma et al.EMNLP 2022 · 57 citations
- FinTextQA: A Dataset for Long-form Financial Question AnsweringJian Chen, Peilin Zhou, Yining Hua, Loh Xin et al.ACL 2024 · 7 citations
Related papers
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu et al.AAAI 2026
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan et al.ACL 2025 · 53 citations
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu et al.CVPR 2025
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb et al.ACL 2025 · 33 citations
- Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsZhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang et al.ACM MM 2025 · 4 citations
