Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking
Haruka Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman, Israel Nir, Ana-Roxana Pop, Nitzan Razin, Sarah Dean, Thorsten Joachims, Udi Weinsberg
摘要
Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker (ESR) generates a candidate set, which is subsequently re-ranked by a late-stage ranker (LSR). While there are many reinforcement learning (RL) methods for training the LSR, end-to-end training of the ESR has proven challenging. In particular, naive application of "vanilla" policy gradient (V-PG) is not scalable for candidate-set sizes relevant for practical use due to exploding variance. This issue arises because V-PG propagates the gradient to the joint probability of the candidate sets, ignoring the contribution of each specific item in the candidate set to the reward. To mitigate this issue, we propose a novel "credit-assigned" policy gradient (CA-PG) , which computes gradients with respect to the probability that the target item is chosen in any candidate set, i.e. marginalizing over all candidate sets that contain it. Our theoretical analysis reveals that CA-PG significantly reduces the variance of V-PG by marginalizing over the specific composition of the candidate set, while preserving the ability to learn the correct ranking of items under a reasonably aligned LSR policy. Experiments on both synthetic and real-world data demonstrate that CA-PG improves the convergence speed and training stability for ESRs utilizing the canonical Plackett-Luce model, especially when the candidate-set size is large.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
- Computationally Efficient Optimization of Plackett-Luce Ranking Models for Relevance and FairnessHarrie OosterhuisSIGIR 2021 · 被引用 68 次
- On Component Interactions in Two-Stage Recommender SystemsJiri Hron, Karl Krauth, Michael I. Jordan, Niki KilbertusNeurIPS 2021 · 被引用 39 次
相关 Paper
- PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation ModelsJeongjae Lee, Jong Chul YeICLR 2026 · 被引用 2 次
- C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented GenerationGuoxin Chen, Minpeng Liao, Peiying Yu, Dingmin Wang 等ICML 2025
- Difference Advantage Estimation for Multi-Agent Policy GradientsYueheng Li, Guangming Xie, Zongqing LuICML 2022 · 被引用 24 次
- Reinforcement Learning to Rank with Pairwise Policy GradientJun Xu, Zeng Wei, Long Xia, Yanyan Lan 等SIGIR 2020 · 被引用 32 次
- Omnia: Efficient RAG Serving through Speculative SchedulingRongtian Fu, Shigang Li, Youxuan Xu, Tong Wu 等HPDC 2026
