Lune

EMNLP2025顶会

OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain

Shuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong Wen

2025年份
6被引次数
8顶会引用

摘要

Retrieval-augmented generation (RAG) has emerged as a key application of large language models (LLMs), especially in vertical domains where LLMs lack domain-specific knowledge. Nevertheless, current RAG benchmarks often suffer from narrow scenarios and limited evaluation dimensions, hindering an all-sides understanding of RAG models in real-world vertical applications. This paper introduces Om-niEval, an omnidirectional and automatic RAG benchmark for the financial domain, featured by its omnidirectional evaluation framework: First, we categorize RAG scenarios by five task classes and 16 financial topics, leading to a matrix-based structured assessment. Next, we leverage a multi-dimensional and auto-chained data generation pipeline that integrates LLMbased automatic generation and human annotation approaches, creating high-quality evaluation instances. Further, we adopt a multi-stage evaluation to assess both retrieval and generation performance, resulting in a holistic RAG evaluation. Finally, rule-based and LLM-based metrics are combined to build a multi-level evaluation system. Our experiments indicate that the performance of RAG systems varies across topics and tasks, highlighting the importance of multi-aspect and structured assessments to better locate the advantages and disadvantages of RAG systems. We release our code at https://github.com/RUC-NLPIR/OmniEval .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖