A Dense Subset Index for Collective Query Coverage
Kartik Nair, Pritish Chakraborty, Atharva Tambat, Indradyumna Roy, Soumen Chakrabarti, Anirban Dasgupta, Abir De
Abstract
In traditional information retrieval, corpus items compete with each other to occupy top ranks in response to a query. In contrast, in many recent retrieval scenarios associated with complex, multi-hop question answering or text-to-SQL, items are not self-complete: they must instead collaborate, i.e., information from multiple items must be combined to respond to the query. In the context of modern dense retrieval, this need translates into finding a small collection of corpus items whose contextual word vectors collectively cover the contextual word vectors of the query. The central challenge is to retrieve a near-optimal collection of covering items in time that is sublinear in corpus size. By establishing coverage as a submodular objective, we enable successive dense index probes to quickly assemble an item collection that achieves near-optimal coverage. Successive query vectors are iteratively `edited', and the dense index is built using random projections of a novel, lifted dense vector space. Beyond rigorous theoretical guarantees, we report on a scalable implementation of this new form of vector database. Extensive experiments establish the empirical success of DISCo, in terms of the best coverage vs. query latency tradeoffs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4ccc8d8-0bc2-47fd-a096-6b4c20bdef51Builds on21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base EmbeddingsApoorv Saxena, Aditay Tripathi, Partha P. TalukdarACL 2020 · 488 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- Generative Retrieval as Multi-Vector Dense RetrievalShiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen et al.SIGIR 2024 · 14 citations
- On Complementarity Objectives for Hybrid RetrievalDohyeon Lee, Seung-won Hwang, Kyungjae Lee, Seungtaek Choi et al.ACL 2023 · 4 citations
- DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational SearchSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasSIGIR 2025 · 5 citations
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang et al.AAAI 2025 · 2 citations
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du et al.ICLR 2021 · 232 citations
