A Dense Subset Index for Collective Query Coverage
Kartik Nair, Pritish Chakraborty, Atharva Tambat, Indradyumna Roy, Soumen Chakrabarti, Anirban Dasgupta, Abir De
摘要
In traditional information retrieval, corpus items compete with each other to occupy top ranks in response to a query. In contrast, in many recent retrieval scenarios associated with complex, multi-hop question answering or text-to-SQL, items are not self-complete: they must instead collaborate, i.e., information from multiple items must be combined to respond to the query. In the context of modern dense retrieval, this need translates into finding a small collection of corpus items whose contextual word vectors collectively cover the contextual word vectors of the query. The central challenge is to retrieve a near-optimal collection of covering items in time that is sublinear in corpus size. By establishing coverage as a submodular objective, we enable successive dense index probes to quickly assemble an item collection that achieves near-optimal coverage. Successive query vectors are iteratively `edited', and the dense index is built using random projections of a novel, lifted dense vector space. Beyond rigorous theoretical guarantees, we report on a scalable implementation of this new form of vector database. Extensive experiments establish the empirical success of DISCo, in terms of the best coverage vs. query latency tradeoffs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsApoorv Saxena, Aditay Tripathi, Partha P. TalukdarACL 2020 · 被引用 488 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
相关 Paper
- Generative Retrieval as Multi-Vector Dense RetrievalShiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen 等SIGIR 2024 · 被引用 14 次
- On Complementarity Objectives for Hybrid RetrievalDohyeon Lee, Seung-won Hwang, Kyungjae Lee, Seungtaek Choi 等ACL 2023 · 被引用 4 次
- DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational SearchSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasSIGIR 2025 · 被引用 5 次
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang 等AAAI 2025 · 被引用 2 次
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du 等ICLR 2021 · 被引用 232 次
