RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora
Hanjun Cho, Jay-Yoon Lee
摘要
Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity. This mismatch undermines evaluation validity: retrievers can be unfairly undervalued even when they retrieve documents that provide sufficient evidence, because redundancy across documents is not accounted for in evaluation. On the other hand, retrievers that perform well on standard benchmarks often generalize poorly to real-world corpora with highly similar and redundant documents. We present RARE (Redundancy-Aware Retrieval Evaluation), a framework for constructing realistic benchmarks by (i) decomposing documents into atomic facts to enable precise redundancy tracking and (ii) enhancing LLM-based data generation with CRRF. RAG benchmark data usually requires multiple quality criteria, but LLMs often yield trivial outputs. CRRF scores criteria separately and fuses decisions by rank, improving the reliability of generated data. Applying RARE to Finance, Legal, and Patent corpora, we introduce RedQA, where a strong retriever baseline drops from 66.4% PerfRecall@10 on 4-hop General-Wiki to 5.0-27.9% PerfRecall@10 at 4-hop depth, revealing robustness gaps that current benchmarks fail to capture. RARE enables practitioners to build domain-specific RAG evaluations that faithfully reflect real-world deployment conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan 等ACL 2025 · 被引用 53 次
- Promptagator: Few-shot Dense Retrieval From 8 ExamplesZhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan 等ICLR 2023 · 被引用 46 次
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai 等EMNLP 2024 · 被引用 8 次
相关 Paper
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb 等ACL 2025 · 被引用 33 次
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question AnsweringTeng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang 等EMNLP 2025 · 被引用 2 次
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu 等AAAI 2026
- Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval EvaluationAndrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti 等ICML 2026 · 被引用 1 次
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 等ICML 2025
