PSLOG: Pretraining with Search Logs for Document Ranking
Zhan Su, Zhicheng Dou, Yujia Zhou, Ziyuan Zhao, Ji-Rong Wen
摘要
Recently, pretrained models have achieved remarkable performance not only in natural language processing but also in information retrieval (IR). Previous studies show that IR-oriented pretraining tasks can achieve better performance than only finetuning pretrained language models in IR datasets. Besides, the massive search log data obtained from mainstream search engines can be used in IR pretraining, for it contains users' implicit judgments of document relevance under a concrete query. However, existing methods mainly use direct query-document click signals to pretrain models. The potential supervision signals from search logs are far from being well explored. In this paper, we propose to comprehensively leverage four query-document relevance relations, including co-interaction and multi-hop relations, to pretrain ranking models in IR. Specifically, we focus on the user's click behavior and construct an Interaction Graph to represent the global relevance relations between queries and documents from all search logs. With the graph, we can consider the co-interaction and multi-hop q-d relationships through their neighbor nodes. Based on the relations extracted from the interaction graph, we propose four strategies to generate contrastive positive and negative q-d pairs and use these data to pretrain ranking models. Experimental results on both industrial and academic datasets demonstrate the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
- A Graph-Enhanced Click Model for Web SearchJianghao Lin, Weiwen Liu, Xinyi Dai, Weinan Zhang 等SIGIR 2021 · 被引用 44 次
- B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc RetrievalXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan 等SIGIR 2021 · 被引用 36 次
- Modeling Intent Graph for Search Result DiversificationZhan Su, Zhicheng Dou, Yutao Zhu, Xubo Qin 等SIGIR 2021 · 被引用 32 次
相关 Paper
- Embedding Prior Task-specific Knowledge into Language Models for Context-aware Document RankingShuting Wang, Yutao Zhu, Zhicheng DouKDD 2025
- Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-RankingTim Baumgärtner, Leonardo F. R. Ribeiro, Nils Reimers, Iryna GurevychEMNLP 2022 · 被引用 3 次
- Distributionally Robust Optimization for Unbiased Learning to RankZechun Niu, Lang Mei, Chong Chen, Jiaxin MaoSIGIR 2025
- Session Search with Pre-trained Graph Classification ModelShengjie Ma, Chong Chen, Jiaxin Mao, Qi Tian 等SIGIR 2023 · 被引用 3 次
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural NetworksYusi Zhang, Chuanjie Liu, Angen Luo, Hui Xue 等WWW 2021 · 被引用 6 次
