Benchmarking Large Language Models in Retrieval-Augmented Generation
Jiawei Chen, Hongyu Lin, Xianpei Han, Le Sun
摘要
Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness. To this end, we establish Retrieval-Augmented Generation Benchmark (RGB), a new corpus for RAG evaluation in both English and Chinese. RGB divides the instances within the benchmark into 4 separate testbeds based on the aforementioned fundamental abilities required to resolve the case. Then we evaluate 6 representative LLMs on RGB to diagnose the challenges of current LLMs when applying RAG. Evaluation reveals that while LLMs exhibit a certain degree of noise robustness, they still struggle significantly in terms of negative rejection, information integration, and dealing with false information. The aforementioned assessment outcomes indicate that there is still a considerable journey ahead to effectively apply RAG to LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper112
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen 等NeurIPS 2025 · 被引用 76 次
- Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam GenerationGauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent CallotICML 2024 · 被引用 35 次
- Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs ReasoningJunming Liu, Siyuan Meng, Yanting Gao, Song Mao 等ICCV 2025 · 被引用 34 次
- CABINET: Content Relevance-based Noise Reduction for Table Question AnsweringSohan Patnaik, Heril Changwal, Milan Aggarwal, Sumit Bhatia 等ICLR 2024 · 被引用 34 次
- MiniCheck: Efficient Fact-Checking of LLMs on Grounding DocumentsLiyan Tang, Philippe Laban, Greg DurrettEMNLP 2024 · 被引用 26 次
它引用的顶会 Paper6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 被引用 187 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
相关 Paper
- Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsJinyang Wu, Shuai Zhang, Feihu Che, Mingkuan Feng 等ACL 2025 · 被引用 12 次
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu 等AAAI 2026
- Re³: Relevance & Recency Retrieval for Mitigating Temporal HallucinationJiawei Cao, Jie Ouyang, Mingyue Cheng, Zhaomeng Zhou 等ACL 2026
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language ModelsCheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu 等ACL 2024
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan 等ACL 2025 · 被引用 53 次
