Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval
Hongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Fuzheng Zhang, Wei Wu
摘要
Recently, the retrieval models based on dense representations have been gradually applied in the first stage of the document retrieval tasks, showing better performance than traditional sparse vector space models. To obtain high efficiency, the basic structure of these models is Bi-encoder in most cases. However, this simple structure may cause serious information loss during the encoding of documents since the queries are agnostic. To address this problem, we design a method to mimic the queries on each of the documents by an iterative clustering process and represent the documents by multiple pseudo queries (i.e., the cluster centroids). To boost the retrieval process using approximate nearest neighbor search library, we also optimize the matching function with a two-step score calculation procedure. Experimental results on several popular ranking and QA datasets show that our model can achieve state-of-the-art results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang 等ACL 2022 · 被引用 80 次
- Generative Retrieval Meets Multi-Graded RelevanceYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等NeurIPS 2024 · 被引用 18 次
- LIDER: An Efficient High-dimensional Learned Index for Large-scale Dense Passage RetrievalYifan Wang, Haodi Ma, Daisy Zhe WangVLDB 2023 · 被引用 17 次
- Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality EstimationYuan Ge, Yilun Liu, Chi Hu, Weibin Meng 等EMNLP 2024 · 被引用 10 次
- VIRT: Improving Representation-based Text Matching via Virtual InteractionDan Li, Yang Yang, Hongyin Tang, Jiahao Liu 等EMNLP 2022 · 被引用 8 次
它引用的顶会 Paper3
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Modularized Transfomer-based Ranking FrameworkLuyu Gao, Zhuyun Dai, Jamie CallanEMNLP 2020 · 被引用 52 次
相关 Paper
- Pseudo-Relevance for Enhancing Document RepresentationJihyuk Kim, Seung-won Hwang, Seoho Song, Hyeseon Ko 等EMNLP 2022 · 被引用 1 次
- Multivariate Representation Learning for Information RetrievalHamed Zamani, Michael BenderskySIGIR 2023 · 被引用 7 次
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang 等AAAI 2025 · 被引用 2 次
- Constructing Tree-based Index for Efficient and Effective Dense RetrievalHaitao Li, Qingyao Ai, Jingtao Zhan, Jiaxin Mao 等SIGIR 2023 · 被引用 21 次
- CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionXingwei He, Yeyun Gong, A-Long Jin, Hang Zhang 等EMNLP 2023 · 被引用 3 次
