RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language Models
Zheng Liu, Shitao Xiao, Yingxia Shao, Zhao Cao
摘要
To better support information retrieval tasks such as web search and open-domain question answering, growing effort is made to develop retrieval-oriented language models, e.g., RetroMAE (Xiao et al., 2022b) and many others (Gao and Callan, 2021; Wang et al., 2021a). Most of the existing works focus on improving the semantic representation capability for the contextualized embedding of the [CLS] token. However, recent study shows that the ordinary tokens besides [CLS] may provide extra information, which help to produce a better representation effect (Lin et al., 2022) . As such, it's necessary to extend the current methods where all contextualized embeddings can be jointly pre-trained for the retrieval tasks. In this work, we propose a novel pre-training method called Duplex Masked Auto-Encoder, a.k.a. DupMAE. It is designed to improve the quality of semantic representation where all contextualized embeddings of the pre-trained model can be leveraged. It takes advantage of two complementary auto-encoding tasks: one reconstructs the input sentence with the [CLS] embedding; the other one predicts the bagof-words feature of the input sentence with the ordinary tokens' embeddings. The two tasks are jointly conducted to train a unified encoder, where the whole contextualized embeddings are aggregated in a compact way to produce the final semantic representation. DupMAE is simple but empirically competitive: it substantially improves the pre-trained model's representation capability and transferability, where superior retrieval performances can be achieved on popular benchmarks, like MS MARCO and BEIR. Our code is released at: https://github.com/staoxiao/RetroMAE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense RetrievalChaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao 等ACL 2024 · 被引用 10 次
- Threshold-driven Pruning with Segmented Maximum Term Weights for Approximate Cluster-based Sparse RetrievalYifan Qiao, Parker Carlson, Shanxiu He, Yingrui Yang 等EMNLP 2024 · 被引用 7 次
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 被引用 5 次
- Training Dense Retrievers with Multiple Positive PassagesBenben Wang, Minghao Tang, Hengran Zhang, Jiafeng Guo 等KDD 2026 · 被引用 1 次
它引用的顶会 Paper13
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
相关 Paper
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 被引用 63 次
- LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale RetrievalTao Shen, Xiubo Geng, Chongyang Tao, Can Xu 等ICLR 2023 · 被引用 14 次
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin 等AAAI 2023 · 被引用 34 次
- Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak DecoderShuqi Lu, Di He, Chenyan Xiong, Guolin Ke 等EMNLP 2021 · 被引用 46 次
- MEXMA: Token-level objectives improve sentence representationsJoão Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc BarraultACL 2025
