Pseudo-Relevance for Enhancing Document Representation
Jihyuk Kim, Seung-won Hwang, Seoho Song, Hyeseon Ko, Young-In Song
Abstract
This paper studies how to enhance the document representation for the bi-encoder approach in dense document retrieval. The bi-encoder, separately encoding a query and a document as a single vector, is favored for high efficiency in large-scale information retrieval, compared to more effective but complex architectures. To combine the strength of the two, the multi-vector representation of documents for bi-encoder, such as ColBERT preserving all token embeddings, has been widely adopted. Our contribution is to reduce the size of the multi-vector representation, without compromising the effectiveness, supervised by query logs. Our proposed solution decreases the latency and the memory footprint, up to 8- and 3-fold, validated on MSMARCO and real-world search query logs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64e89921-9059-4165-8bf6-5643849a56c6Builds on6
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Intra-Document Cascading: Learning to Select Passages for Neural Document RankingSebastian Hofstätter, Bhaskar Mitra, Hamed Zamani, Nick Craswell et al.SIGIR 2021 · 35 citations
Related papers
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang et al.ACL 2021
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ACL 2022 · 80 citations
- CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction RetrievalRohit Kumar Salla, Manoj Saravanan, Ramya AmancherlaICML 2026
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector RetrievalLixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng et al.ICML 2026
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei et al.NeurIPS 2023 · 60 citations
