Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval
Luyu Gao, Jamie Callan
摘要
Recent research demonstrates the effectiveness of using fine-tuned language models (LM) for dense retrieval. However, dense retrievers are hard to train, typically requiring heavily engineered fine-tuning pipelines to realize their full potential. In this paper, we identify and address two underlying problems of dense retrievers: i) fragility to training data noise and ii) requiring large batches to robustly learn the embedding space. We use the recently proposed Condenser pre-training architecture, which learns to condense information into the dense vector through LM pre-training. On top of it, we propose coCondenser, which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space. Retrieval experiments on MS-MARCO, Natural Question, and Trivia QA datasets show that coCondenser removes the need for heavy data engineering such as augmentation, synthesis, or filtering, as well as the need for large batch training. It shows comparable performance to RocketQA, a state-of-the-art, heavily engineered system, using simple small batch finetuning. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper76
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao 等NeurIPS 2022 · 被引用 242 次
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 被引用 211 次
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv 等ICLR 2022 · 被引用 137 次
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong 等SIGIR 2023 · 被引用 68 次
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 被引用 63 次
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
相关 Paper
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao 等ACL 2023 · 被引用 41 次
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin 等AAAI 2023 · 被引用 34 次
- A Gradient Accumulation Method for Dense Retriever under Memory ConstraintJaehee Kim, Yukyung Lee, Pilsung KangNeurIPS 2024 · 被引用 10 次
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversXueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin 等ACL 2025 · 被引用 20 次
