CAPSTONE: Curriculum Sampling for Dense Retrieval with Document Expansion
Xingwei He, Yeyun Gong, A-Long Jin, Hang Zhang, Anlei Dong, Jian Jiao, Siu-Ming Yiu, Nan Duan
摘要
The dual-encoder has become the de facto architecture for dense retrieval. Typically, it computes the latent representations of the query and document independently, thus failing to fully capture the interactions between the query and document. To alleviate this, recent research has focused on obtaining query-informed document representations. During training, it expands the document with a real query, but during inference, it replaces the real query with a generated one. This inconsistency between training and inference causes the dense retrieval model to prioritize query information while disregarding the document when computing the document representation. Consequently, it performs even worse than the vanilla dense retrieval model because its performance heavily relies on the relevance between the generated queries and the real query. In this paper, we propose a curriculum sampling strategy that utilizes pseudo queries during training and progressively enhances the relevance between the generated query and the real query. By doing so, the retrieval model learns to extend its attention from the document alone to both the document and query, resulting in high-quality query-informed document representations. Experimental results on both in-domain and out-of-domain datasets demonstrate that our approach outperforms previous dense retrieval models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng 等AAAI 2020 · 被引用 885 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
相关 Paper
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin 等EMNLP 2023 · 被引用 2 次
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong 等EMNLP 2022 · 被引用 6 次
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang 等ACL 2021
- Contextualized Query Embeddings for Conversational SearchSheng-Chieh Lin, Jheng-Hong Yang, Jimmy LinEMNLP 2021 · 被引用 40 次
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv 等ICLR 2022 · 被引用 137 次
