Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation
Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Denton, Yuan Xue
摘要
Retrieval quality is the primary bottleneck for accuracy and robustness in retrieval-augmented generation (RAG). Current evaluation relies on heuristically constructed query sets, which introduce a hidden intrinsic bias. We formalize retrieval evaluation as a statistical estimation problem, showing that metric reliability is fundamentally limited by the evaluation-set construction. We further introduce semantic stratification, which grounds evaluation in corpus structure by organizing documents into an interpretable global space of entity-based clusters and systematically generating queries for missing strata. This yields (1) formal semantic coverage guarantees across retrieval regimes and (2) interpretable visibility into retrieval failure modes. Experiments across multiple benchmarks and retrieval methods validate our framework. The results expose systematic coverage gaps, identify structural signals that explain variance in retrieval performance, and show that stratified evaluation yields more stable and transparent assessments while supporting more trustworthy decision-making than aggregate metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan 等ACL 2025 · 被引用 53 次
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey 等ACL 2020 · 被引用 20 次
相关 Paper
- Testing Retrieval-Augmented Generation Systems with Chunk CoverageJinhan Kim, Samuele Pasini, Paolo TonellaISSTA 2026
- SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity ReductionLu Dai, Yijie Xu, Jinhui Ye, Hao Liu 等ICLR 2025
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity CorporaHanjun Cho, Jay-Yoon LeeACL 2026 · 被引用 1 次
- Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationZhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das 等EMNLP 2024 · 被引用 2 次
- Unanswerability Evaluation for Retrieval Augmented GenerationXiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng WuACL 2025 · 被引用 8 次
