A Unified Pretraining Framework for Passage Ranking and Expansion
Ming Yan, Chenliang Li, Bin Bi, Wei Wang, Songfang Huang
摘要
Pretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Perturbation-Invariant Adversarial Training for Neural Ranking Models: Improving the Effectiveness-Robustness Trade-OffYu-An Liu, Ruqing Zhang, Mingkun Zhang, Wei Chen 等AAAI 2024 · 被引用 17 次
- Incorporating Explicit Knowledge in Pre-trained Language Models for Passage Re-rankingQian Dong, Yiding Liu, Suqi Cheng, Shuaiqiang Wang 等SIGIR 2022 · 被引用 14 次
- Attack-in-the-Chain: Bootstrapping Large Language Models for Attacks Against Black-Box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等AAAI 2025 · 被引用 11 次
- Stop Hardening Everything: A Training-Free Neuron-Level Defense for Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Hongru Song, Jiafeng Guo 等ACL 2026
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan 等EMNLP 2020 · 被引用 41 次
相关 Paper
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan 等EMNLP 2022 · 被引用 69 次
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin 等AAAI 2023 · 被引用 34 次
- List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented GenerationShicheng Xu, Liang Pang, Jun Xu, Huawei Shen 等WWW 2024 · 被引用 13 次
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM GuidanceSijia Yao, Pengcheng Huang, Zhenghao Liu, Yu Gu 等EMNLP 2025 · 被引用 6 次
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao 等EMNLP 2021 · 被引用 147 次
