Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval
Hongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Fuzheng Zhang, Wei Wu
Abstract
Recently, the retrieval models based on dense representations have been gradually applied in the first stage of the document retrieval tasks, showing better performance than traditional sparse vector space models. To obtain high efficiency, the basic structure of these models is Bi-encoder in most cases. However, this simple structure may cause serious information loss during the encoding of documents since the queries are agnostic. To address this problem, we design a method to mimic the queries on each of the documents by an iterative clustering process and represent the documents by multiple pseudo queries (i.e., the cluster centroids). To boost the retrieval process using approximate nearest neighbor search library, we also optimize the matching function with a two-step score calculation procedure. Experimental results on several popular ranking and QA datasets show that our model can achieve state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dedd056-43b4-4985-b70f-bfd804f52a59Cited by top-tier papers11
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ACL 2022 · 80 citations
- Generative Retrieval Meets Multi-Graded RelevanceYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.NeurIPS 2024 · 18 citations
- LIDER: An Efficient High-dimensional Learned Index for Large-scale Dense Passage RetrievalYifan Wang, Haodi Ma, Daisy Zhe WangVLDB 2023 · 17 citations
- Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality EstimationYuan Ge, Yilun Liu, Chi Hu, Weibin Meng et al.EMNLP 2024 · 10 citations
- VIRT: Improving Representation-based Text Matching via Virtual InteractionDan Li, Yang Yang, Hongyin Tang, Jiahao Liu et al.EMNLP 2022 · 8 citations
Builds on3
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Modularized Transfomer-based Ranking FrameworkLuyu Gao, Zhuyun Dai, Jamie CallanEMNLP 2020 · 52 citations
Related papers
- Pseudo-Relevance for Enhancing Document RepresentationJihyuk Kim, Seung-won Hwang, Seoho Song, Hyeseon Ko et al.EMNLP 2022 · 1 citation
- Multivariate Representation Learning for Information RetrievalHamed Zamani, Michael BenderskySIGIR 2023 · 7 citations
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang et al.AAAI 2025 · 2 citations
- Constructing Tree-based Index for Efficient and Effective Dense RetrievalHaitao Li, Qingyao Ai, Jingtao Zhan, Jiaxin Mao et al.SIGIR 2023 · 21 citations
- CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionXingwei He, Yeyun Gong, A-Long Jin, Hang Zhang et al.EMNLP 2023 · 3 citations
