Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation
Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Denton, Yuan Xue
Abstract
Retrieval quality is the primary bottleneck for accuracy and robustness in retrieval-augmented generation (RAG). Current evaluation relies on heuristically constructed query sets, which introduce a hidden intrinsic bias. We formalize retrieval evaluation as a statistical estimation problem, showing that metric reliability is fundamentally limited by the evaluation-set construction. We further introduce semantic stratification, which grounds evaluation in corpus structure by organizing documents into an interpretable global space of entity-based clusters and systematically generating queries for missing strata. This yields (1) formal semantic coverage guarantees across retrieval regimes and (2) interpretable visibility into retrieval failure modes. Experiments across multiple benchmarks and retrieval methods validate our framework. The results expose systematic coverage gaps, identify structural signals that explain variance in retrieval performance, and show that stratified evaluation yields more stable and transparent assessments while supporting more trustworthy decision-making than aggregate metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca60a105-e11d-4603-a80c-f8f059391136Builds on3
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan et al.ACL 2025 · 53 citations
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey et al.ACL 2020 · 20 citations
Related papers
- Testing Retrieval-Augmented Generation Systems with Chunk CoverageJinhan Kim, Samuele Pasini, Paolo TonellaISSTA 2026
- SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity ReductionLu Dai, Yijie Xu, Jinhui Ye, Hao Liu et al.ICLR 2025
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity CorporaHanjun Cho, Jay-Yoon LeeACL 2026 · 1 citation
- Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationZhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das et al.EMNLP 2024 · 2 citations
- Unanswerability Evaluation for Retrieval Augmented GenerationXiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng WuACL 2025 · 8 citations
