Reinforcement Learning to Rank with Pairwise Policy Gradient
Jun Xu, Zeng Wei, Long Xia, Yanyan Lan, Dawei Yin, Xueqi Cheng, Ji-Rong Wen
Abstract
This paper concerns reinforcement learning (RL) of the document ranking models for information retrieval (IR). One branch of the RL approaches to ranking formalize the process of ranking with Markov decision process (MDP) and determine the model parameters with policy gradient. Though preliminary success has been shown, these approaches are still far from achieving their full potentials. Existing policy gradient methods directly utilize the absolute performance scores (returns) of the sampled document lists in its gradient estimations, which may cause two limitations: 1) fail to reflect the relative goodness of documents within the same query, which usually is close to the nature of IR ranking; 2) generate high variance gradient estimations, resulting in slow learning speed and low ranking accuracy. To deal with the issues, we propose a novel policy gradient algorithm in which the gradients are determined using pairwise comparisons of two document lists sampled within the same query. The algorithm, referred to as Pairwise Policy Gradient (PPG), repeatedly samples pairs of document lists, estimates the gradients with pairwise comparisons, and finally updates the model parameters. Theoretical analysis shows that PPG makes an unbiased and low variance gradient estimations. Experimental results have demonstrated performance gains over the state-of-the-art baselines in search result diversification and text retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eebac6dc-d3fa-4c2d-a9fd-9d43fc5d0a1dCited by top-tier papers6
- Knowledge Enhanced Search Result DiversificationZhan Su, Zhicheng Dou, Yutao Zhu, Ji-Rong WenKDD 2022 · 16 citations
- Optimize What You Evaluate With: Search Result Diversification Based on Metric OptimizationHai-Tao YuAAAI 2022 · 11 citations
- Unified Off-Policy Learning to Rank: a Reinforcement Learning PerspectiveZeyu Zhang, Yi Su, Hui Yuan, Yiran Wu et al.NeurIPS 2023 · 9 citations
- Bridging the Preference Gap between Retrievers and LLMsZixuan Ke, Weize Kong, Cheng Li, Mingyang Zhang et al.ACL 2024 · 8 citations
- MA4DIV: Multi-Agent Reinforcement Learning for Search Result DiversificationYiqun Chen, Jiaxin Mao, Yi Zhang, Dehong Ma et al.WWW 2025 · 7 citations
Related papers
- Allowing for The Grounded Use of Temporal Difference Learning in Large Ranking Models via Substate UpdatesDaniel CohenSIGIR 2021 · 1 citation
- Ranking Policy GradientKaixiang Lin, Jiayu ZhouICLR 2020 · 8 citations
- Topic-oriented Adversarial Attacks against Black-box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2023 · 20 citations
- Lightweight and Direct Document Relevance Optimization for Generative Information RetrievalKidist Amde Mekonnen, Yubao Tang, Maarten de RijkeSIGIR 2025 · 3 citations
- PairDistill: Pairwise Relevance Distillation for Dense RetrievalChao-Wei Huang, Yun-Nung ChenEMNLP 2024 · 3 citations
