NuggetIndex: Governed Atomic Retrieval for Maintainable RAG
Saber Zerhoudi, Michael Granitzer, Jelena Mitrovic
Abstract
Retrieval-augmented generation (RAG) systems are frequently evaluated via fact-based metrics, yet standard implementations retrieve passages or static propositions. This unit mismatch between evaluation and retrieval objects hinders maintenance when corpora evolve and fails to capture superseded facts or source disagreements. We propose NuggetIndex, a retrieval system that stores atomic information units as managed records, so called nuggets. Each record maintains links to evidence, a temporal validity interval, and a lifecycle state. By filtering invalid or deprecated nuggets prior to ranking, the system prevents the inclusion of outdated information. We evaluate the approach using a nuggetized MS MARCO subset, a temporal Wikipedia QA dataset, and a multi-hop QA task. Against passage and unmanaged proposition retrieval baselines, NuggetIndex improves nugget recall by 42%, increases temporal correctness by 9 percentage points without the recall collapse observed in time-filtered baselines, and reduces conflict rates by 55%. The compact nugget format reduces generator input length by 64% while enabling lightweight index structures suitable for browser-based and resource-constrained deployment. We release our implementation, datasets, and evaluation scripts
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Dense X Retrieval: What Retrieval Granularity Should We Use?Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu et al.EMNLP 2024 · 52 citations
- SituatedQA: Incorporating Extra-Linguistic Contexts into QAMichael J. Q. Zhang, Eunsol ChoiEMNLP 2021 · 2 citations
Related papers
- The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language ModelsRonak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos et al.SIGIR 2025 · 13 citations
- NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence ChainsShiyao Peng, Qianhe Zheng, Zhuodi Hao, Zichen Tang et al.WWW 2026
- Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval EvaluationAndrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti et al.ICML 2026 · 1 citation
- HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented GenerationJie Ouyang, Tingyue Pan, Mingyue Cheng, Ruiran Yan et al.ACL 2025 · 14 citations
- Re³: Relevance & Recency Retrieval for Mitigating Temporal HallucinationJiawei Cao, Jie Ouyang, Mingyue Cheng, Zhaomeng Zhou et al.ACL 2026
