Testing Retrieval-Augmented Generation Systems with Chunk Coverage
Jinhan Kim, Samuele Pasini, Paolo Tonella
Abstract
Retrieval-Augmented Generation (RAG)-based systems 1 are increasingly deployed in high-stakes settings where correct behaviour depends not only on the language model but also on the retrieval component that selects external documents at inference time. While existing RAG evaluation metrics assess retrieval and generation quality on a per-query basis, typically relying on query-level test oracles such as reference answers or relevance annotations, they provide limited insight into whether a test suite adequately exercises the retrieval behaviour of the system as a whole. In this paper, we introduce Chunk Coverage (CC), an oracleindependent test adequacy criterion for testing the retrieval component of RAG systems. CC measures the fraction of corpus chunks that are retrieved at least once across a test suite, providing a structural view of which parts of the retrieval space have been exercised. We further show how CC can be used to guide test selection and generation by prioritising queries that expand coverage of previously unexercised retrieval regions. We evaluate CC on clinical and financial RAG system scenarios. CC-guided testing reaches 50% of attainable coverage 1.7× faster than random selection and 4.2× faster than redundancy-biased strategies. Moreover, CC improves fault detection effectiveness (APFD) by 10% to 25% over random, indicating earlier discovery of distinct retrieval faults. These results show that CC captures retrieval diversity relevant to effective testing without requiring test oracles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90664f36-2c0f-48ad-9480-eb008a110ae6Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
Related papers
- Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval EvaluationAndrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti et al.ICML 2026 · 1 citation
- A New HOPE: Domain-agnostic Automatic Evaluation of Text ChunkingHenrik Brådland, Morten Goodwin, Per-Arne Andersen, Alexander Salveson Nossum et al.SIGIR 2025 · 9 citations
- Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationZhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das et al.EMNLP 2024 · 2 citations
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity CorporaHanjun Cho, Jay-Yoon LeeACL 2026 · 1 citation
- Grounding Language Model with Chunking-Free In-Context RetrievalHongjin Qian, Zheng Liu, Kelong Mao, Yujia Zhou et al.ACL 2024
