A Unified Pretraining Framework for Passage Ranking and Expansion
Ming Yan, Chenliang Li, Bin Bi, Wei Wang, Songfang Huang
Abstract
Pretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext deed39d3-0859-4dec-baee-70d519a0403eCited by top-tier papers4
- Perturbation-Invariant Adversarial Training for Neural Ranking Models: Improving the Effectiveness-Robustness Trade-OffYu-An Liu, Ruqing Zhang, Mingkun Zhang, Wei Chen et al.AAAI 2024 · 17 citations
- Incorporating Explicit Knowledge in Pre-trained Language Models for Passage Re-rankingQian Dong, Yiding Liu, Suqi Cheng, Shuaiqiang Wang et al.SIGIR 2022 · 14 citations
- Attack-in-the-Chain: Bootstrapping Large Language Models for Attacks Against Black-Box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.AAAI 2025 · 11 citations
- Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Hongru Song, Jiafeng Guo et al.ACL 2026
Builds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan et al.EMNLP 2020 · 41 citations
Related papers
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin et al.AAAI 2023 · 34 citations
- List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented GenerationShicheng Xu, Liang Pang, Jun Xu, Huawei Shen et al.WWW 2024 · 13 citations
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM GuidanceSijia Yao, Pengcheng Huang, Zhenghao Liu, Yu Gu et al.EMNLP 2025 · 6 citations
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao et al.EMNLP 2021 · 147 citations
