SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei
Abstract
In this paper, we propose SIMLM (Similarity matching with Language Model pre-training), a simple yet effective pre-training method for dense passage retrieval. It employs a simple bottleneck architecture that learns to compress the passage information into a dense vector through self-supervised pre-training. We use a replaced language modeling objective, which is inspired by ELECTRA (Clark et al., 2020), to improve the sample efficiency and reduce the mismatch of the input distribution between pre-training and fine-tuning. SIMLM only requires access to an unlabeled corpus and is more broadly applicable when there are no labeled data or queries. We conduct experiments on several large-scale passage retrieval datasets and show substantial improvements over strong baselines under various settings. Remarkably, SIMLM even outperforms multivector approaches such as ColBERTv2 (Santhanam et al., 2021) which incurs significantly more storage cost. Our code and model checkpoints are available at https://github.com/ microsoft/unilm/tree/master/simlm .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3205792-1177-4fe3-8ed0-4c3e02d606e8Cited by top-tier papers39
- Learning to Rank in Generative RetrievalYongqi Li, Nan Yang, Liang Wang, Furu Wei et al.AAAI 2024 · 83 citations
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong et al.SIGIR 2023 · 68 citations
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 60 citations
- Fine-Grained Distillation for Long Document RetrievalYucheng Zhou, Tao Shen, Xiubo Geng, Chongyang Tao et al.AAAI 2024 · 44 citations
- Multiview Identifiers Enhanced Generative RetrievalYongqi Li, Nan Yang, Liang Wang, Furu Wei et al.ACL 2023 · 30 citations
Builds on12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
Related papers
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin et al.AAAI 2023 · 34 citations
- Unsupervised Corpus Aware Language Model Pre-training for Dense Passage RetrievalLuyu Gao, Jamie CallanACL 2022
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
- Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense RetrievalChaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao et al.ACL 2024 · 10 citations
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
