Reinforcement Learning to Rank with Pairwise Policy Gradient
Jun Xu, Zeng Wei, Long Xia, Yanyan Lan, Dawei Yin, Xueqi Cheng, Ji-Rong Wen
摘要
This paper concerns reinforcement learning (RL) of the document ranking models for information retrieval (IR). One branch of the RL approaches to ranking formalize the process of ranking with Markov decision process (MDP) and determine the model parameters with policy gradient. Though preliminary success has been shown, these approaches are still far from achieving their full potentials. Existing policy gradient methods directly utilize the absolute performance scores (returns) of the sampled document lists in its gradient estimations, which may cause two limitations: 1) fail to reflect the relative goodness of documents within the same query, which usually is close to the nature of IR ranking; 2) generate high variance gradient estimations, resulting in slow learning speed and low ranking accuracy. To deal with the issues, we propose a novel policy gradient algorithm in which the gradients are determined using pairwise comparisons of two document lists sampled within the same query. The algorithm, referred to as Pairwise Policy Gradient (PPG), repeatedly samples pairs of document lists, estimates the gradients with pairwise comparisons, and finally updates the model parameters. Theoretical analysis shows that PPG makes an unbiased and low variance gradient estimations. Experimental results have demonstrated performance gains over the state-of-the-art baselines in search result diversification and text retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Knowledge Enhanced Search Result DiversificationZhan Su, Zhicheng Dou, Yutao Zhu, Ji-Rong WenKDD 2022 · 被引用 16 次
- Optimize What You Evaluate With: Search Result Diversification Based on Metric OptimizationHai-Tao YuAAAI 2022 · 被引用 11 次
- Unified Off-Policy Learning to Rank: a Reinforcement Learning PerspectiveZeyu Zhang, Yi Su, Hui Yuan, Yiran Wu 等NeurIPS 2023 · 被引用 9 次
- Bridging the Preference Gap between Retrievers and LLMsZixuan Ke, Weize Kong, Cheng Li, Mingyang Zhang 等ACL 2024 · 被引用 8 次
- MA4DIV: Multi-Agent Reinforcement Learning for Search Result DiversificationYiqun Chen, Jiaxin Mao, Yi Zhang, Dehong Ma 等WWW 2025 · 被引用 7 次
相关 Paper
- Allowing for The Grounded Use of Temporal Difference Learning in Large Ranking Models via Substate UpdatesDaniel CohenSIGIR 2021 · 被引用 1 次
- Ranking Policy GradientKaixiang Lin, Jiayu ZhouICLR 2020 · 被引用 8 次
- Topic-oriented Adversarial Attacks against Black-box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等SIGIR 2023 · 被引用 20 次
- Lightweight and Direct Document Relevance Optimization for Generative Information RetrievalKidist Amde Mekonnen, Yubao Tang, Maarten de RijkeSIGIR 2025 · 被引用 3 次
- PairDistill: Pairwise Relevance Distillation for Dense RetrievalChao-Wei Huang, Yun-Nung ChenEMNLP 2024 · 被引用 3 次
