Testing Retrieval-Augmented Generation Systems with Chunk Coverage
Jinhan Kim, Samuele Pasini, Paolo Tonella
摘要
Retrieval-Augmented Generation (RAG)-based systems 1 are increasingly deployed in high-stakes settings where correct behaviour depends not only on the language model but also on the retrieval component that selects external documents at inference time. While existing RAG evaluation metrics assess retrieval and generation quality on a per-query basis, typically relying on query-level test oracles such as reference answers or relevance annotations, they provide limited insight into whether a test suite adequately exercises the retrieval behaviour of the system as a whole. In this paper, we introduce Chunk Coverage (CC), an oracleindependent test adequacy criterion for testing the retrieval component of RAG systems. CC measures the fraction of corpus chunks that are retrieved at least once across a test suite, providing a structural view of which parts of the retrieval space have been exercised. We further show how CC can be used to guide test selection and generation by prioritising queries that expand coverage of previously unexercised retrieval regions. We evaluate CC on clinical and financial RAG system scenarios. CC-guided testing reaches 50% of attainable coverage 1.7× faster than random selection and 4.2× faster than redundancy-biased strategies. Moreover, CC improves fault detection effectiveness (APFD) by 10% to 25% over random, indicating earlier discovery of distinct retrieval faults. These results show that CC captures retrieval diversity relevant to effective testing without requiring test oracles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 被引用 99 次
相关 Paper
- Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval EvaluationAndrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti 等ICML 2026 · 被引用 1 次
- A New HOPE: Domain-agnostic Automatic Evaluation of Text ChunkingHenrik Brådland, Morten Goodwin, Per-Arne Andersen, Alexander Salveson Nossum 等SIGIR 2025 · 被引用 9 次
- Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationZhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das 等EMNLP 2024 · 被引用 2 次
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity CorporaHanjun Cho, Jay-Yoon LeeACL 2026 · 被引用 1 次
- Grounding Language Model with Chunking-Free In-Context RetrievalHongjin Qian, Zheng Liu, Kelong Mao, Yujia Zhou 等ACL 2024
