ConTextual Masked Auto-Encoder for Dense Passage Retrieval
Xing Wu, Guangyuan Ma, Meng Lin, Zijia Lin, Zhongyuan Wang, Songlin Hu
Abstract
Dense passage retrieval aims to retrieve the relevant passages of a query from a large corpus based on dense representations (i.e., vectors) of the query and the passages. Recent studies have explored improving pre-trained language models to boost dense retrieval performance. This paper proposes CoT-MAE (ConTextual Masked Auto-Encoder), a simple yet effective generative pre-training method for dense passage retrieval. CoT-MAE employs an asymmetric encoder-decoder architecture that learns to compress the sentence semantics into a dense vector through self-supervised and context-supervised masked auto-encoding. Precisely, self-supervised masked auto-encoding learns to model the semantics of the tokens inside a text span, and context-supervised masked auto-encoding learns to model the semantical correlation between the text spans. We conduct experiments on large-scale passage retrieval benchmarks and show considerable improvements over strong baselines, demonstrating the high efficiency of CoT-MAE. Our code is available at https://github.com/caskcsg/ir/tree/main/cotmae.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edb76980-7b6c-4fe3-a351-f2f5a50202a8Cited by top-tier papers11
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong et al.SIGIR 2023 · 68 citations
- Task-level Distributionally Robust Optimization for Large Language Model-based Dense RetrievalGuangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su et al.AAAI 2025 · 6 citations
- Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence RegularizationShiqi Wang, Yeqin Zhang, Cam-Tu NguyenAAAI 2024 · 6 citations
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 5 citations
- DELTA: Pre-Train a Discriminative Encoder for Legal Case Retrieval via Structural Word AlignmentHaitao Li, Qingyao Ai, Xinyan Han, Jia Chen et al.AAAI 2025 · 3 citations
Builds on17
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
Related papers
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 63 citations
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
- SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao et al.ACL 2023 · 41 citations
- Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak DecoderShuqi Lu, Di He, Chenyan Xiong, Guolin Ke et al.EMNLP 2021 · 46 citations
- Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span PredictionXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan et al.SIGIR 2022 · 33 citations
