Dense X Retrieval: What Retrieval Granularity Should We Use?
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, Dong Yu
Abstract
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks. When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g. document, passage, or sentence. We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks. Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval. Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format. We conduct an empirical comparison of different retrieval granularity. Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks. Moreover, constructing prompts with fine-grained retrieved units for retrieval-augmented language models improves the performance of downstream QA tasks given a specific computation budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- GFM-RAG: Graph Foundation Model for Retrieval Augmented GenerationLinhao Luo, Zicheng Zhao, Reza Haffari, Dinh Phung et al.NeurIPS 2025 · 54 citations
- AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaJiaxin Bai, Wei Fan, Qi Hu, Qing Zong et al.ACL 2026 · 27 citations
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and GranularitiesWoongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek et al.ACL 2026 · 14 citations
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented GenerationBaolei Zhang, Haoran Xin, Yuxi Chen, Zhuqing Liu et al.S&P 2026 · 12 citations
- Leveraging Attention to Effectively Compress Prompts for Long-Context LLMsYunlong Zhao, Haoran Wu, Bo XuAAAI 2025 · 10 citations
Builds on18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- Phrase Retrieval Learns Passage Retrieval, TooJinhyuk Lee, Alexander Wettig, Danqi ChenEMNLP 2021 · 1 citation
- Improve Dense Passage Retrieval with Entailment TuningLu Dai, Hao Liu, Hui XiongEMNLP 2024
- NuggetIndex: Governed Atomic Retrieval for Maintainable RAGSaber Zerhoudi, Michael Granitzer, Jelena MitrovicSIGIR 2026
- REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question AnsweringYuhao Wang, Ruiyang Ren, Junyi Li, Xin Zhao et al.EMNLP 2024 · 12 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
