PSLOG: Pretraining with Search Logs for Document Ranking
Zhan Su, Zhicheng Dou, Yujia Zhou, Ziyuan Zhao, Ji-Rong Wen
Abstract
Recently, pretrained models have achieved remarkable performance not only in natural language processing but also in information retrieval (IR). Previous studies show that IR-oriented pretraining tasks can achieve better performance than only finetuning pretrained language models in IR datasets. Besides, the massive search log data obtained from mainstream search engines can be used in IR pretraining, for it contains users' implicit judgments of document relevance under a concrete query. However, existing methods mainly use direct query-document click signals to pretrain models. The potential supervision signals from search logs are far from being well explored. In this paper, we propose to comprehensively leverage four query-document relevance relations, including co-interaction and multi-hop relations, to pretrain ranking models in IR. Specifically, we focus on the user's click behavior and construct an Interaction Graph to represent the global relevance relations between queries and documents from all search logs. With the graph, we can consider the co-interaction and multi-hop q-d relationships through their neighbor nodes. Based on the relations extracted from the interaction graph, we propose four strategies to generate contrastive positive and negative q-d pairs and use these data to pretrain ranking models. Experimental results on both industrial and academic datasets demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 512eaa2f-c9d8-4c91-a11a-405d94437308Builds on8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
- A Graph-Enhanced Click Model for Web SearchJianghao Lin, Weiwen Liu, Xinyi Dai, Weinan Zhang et al.SIGIR 2021 · 44 citations
- B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc RetrievalXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan et al.SIGIR 2021 · 36 citations
- Modeling Intent Graph for Search Result DiversificationZhan Su, Zhicheng Dou, Yutao Zhu, Xubo Qin et al.SIGIR 2021 · 32 citations
Related papers
- Embedding Prior Task-specific Knowledge into Language Models for Context-aware Document RankingShuting Wang, Yutao Zhu, Zhicheng DouKDD 2025
- Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-RankingTim Baumgärtner, Leonardo F. R. Ribeiro, Nils Reimers, Iryna GurevychEMNLP 2022 · 3 citations
- Distributionally Robust Optimization for Unbiased Learning to RankZechun Niu, Lang Mei, Chong Chen, Jiaxin MaoSIGIR 2025
- Session Search with Pre-trained Graph Classification ModelShengjie Ma, Chong Chen, Jiaxin Mao, Qi Tian et al.SIGIR 2023 · 3 citations
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural NetworksYusi Zhang, Chuanjie Liu, Angen Luo, Hui Xue et al.WWW 2021 · 6 citations
