RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder
Shitao Xiao, Zheng Liu, Yingxia Shao, Zhao Cao
Abstract
Despite pre-training’s progress in many important NLP tasks, it remains to explore effective pre-training strategies for dense retrieval. In this paper, we propose RetroMAE, a new retrieval oriented pre-training paradigm based on Masked Auto-Encoder (MAE). RetroMAE is highlighted by three critical designs. 1) A novel MAE workflow, where the input sentence is polluted for encoder and decoder with different masks. The sentence embedding is generated from the encoder’s masked input; then, the original sentence is recovered based on the sentence embedding and the decoder’s masked input via masked language modeling. 2) Asymmetric model structure, with a full-scale BERT like transformer as encoder, and a one-layer transformer as decoder. 3) Asymmetric masking ratios, with a moderate ratio for encoder: 15 30%, and an aggressive ratio for decoder: 50 70%. Our framework is simple to realize and empirically competitive: the pre-trained models dramatically improve the SOTA performances on a wide range of dense retrieval benchmarks, like BEIR and MS MARCO. The source code and pre-trained models are made publicly available at https://github.com/staoxiao/RetroMAE so as to inspire more interesting research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a52e6de3-92dd-4ec5-989b-ad1b72ffb19cCited by top-tier papers42
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 60 citations
- Lexically-Accelerated Dense RetrievalHrishikesh Kulkarni, Sean MacAvaney, Nazli Goharian, Ophir FriederSIGIR 2023 · 30 citations
- Scaling Laws For Dense RetrievalYan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao et al.SIGIR 2024 · 26 citations
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversXueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin et al.ACL 2025 · 20 citations
- Generative Retrieval via Term Set GenerationPeitian Zhang, Zheng Liu, Yujia Zhou, Zhicheng Dou et al.SIGIR 2024 · 13 citations
Builds on12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
Related papers
- RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language ModelsZheng Liu, Shitao Xiao, Yingxia Shao, Zhao CaoACL 2023 · 7 citations
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin et al.AAAI 2023 · 34 citations
- LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale RetrievalTao Shen, Xiubo Geng, Chongyang Tao, Can Xu et al.ICLR 2023 · 14 citations
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 5 citations
- Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak DecoderShuqi Lu, Di He, Chenyan Xiong, Guolin Ke et al.EMNLP 2021 · 46 citations
