Dense X Retrieval: What Retrieval Granularity Should We Use?
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, Dong Yu
摘要
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks. When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g. document, passage, or sentence. We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks. Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval. Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format. We conduct an empirical comparison of different retrieval granularity. Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks. Moreover, constructing prompts with fine-grained retrieved units for retrieval-augmented language models improves the performance of downstream QA tasks given a specific computation budget.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- GFM-RAG: Graph Foundation Model for Retrieval Augmented GenerationLinhao Luo, Zicheng Zhao, Reza Haffari, Dinh Phung 等NeurIPS 2025 · 被引用 54 次
- AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaJiaxin Bai, Wei Fan, Qi Hu, Qing Zong 等ACL 2026 · 被引用 27 次
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and GranularitiesWoongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek 等ACL 2026 · 被引用 14 次
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented GenerationBaolei Zhang, Haoran Xin, Yuxi Chen, Zhuqing Liu 等S&P 2026 · 被引用 12 次
- Leveraging Attention to Effectively Compress Prompts for Long-Context LLMsYunlong Zhao, Haoran Wu, Bo XuAAAI 2025 · 被引用 10 次
它引用的顶会 Paper18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
相关 Paper
- Phrase Retrieval Learns Passage Retrieval, TooJinhyuk Lee, Alexander Wettig, Danqi ChenEMNLP 2021 · 被引用 1 次
- Improve Dense Passage Retrieval with Entailment TuningLu Dai, Hao Liu, Hui XiongEMNLP 2024
- NuggetIndex: Governed Atomic Retrieval for Maintainable RAGSaber Zerhoudi, Michael Granitzer, Jelena MitrovicSIGIR 2026
- REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question AnsweringYuhao Wang, Ruiyang Ren, Junyi Li, Xin Zhao 等EMNLP 2024 · 被引用 12 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
