Ranking Policy Gradient
Kaixiang Lin, Jiayu Zhou
Abstract
Sample inefficiency is a long-lasting problem in reinforcement learning (RL). The state-of-the-art estimates the optimal action values while it usually involves an extensive search over the state-action space and unstable optimization. Towards the sample-efficient RL, we propose ranking policy gradient (RPG), a policy gradient method that learns the optimal rank of a set of discrete actions. To accelerate the learning of policy gradient methods, we establish the equivalence between maximizing the lower bound of return and imitating a near-optimal policy without accessing any oracles. These results lead to a general off-policy learning framework, which preserves the optimality, reduces variance, and improves the sample-efficiency. Furthermore, the sample complexity of RPG does not depend on the dimension of state space, which enables RPG for large-scale problems. We conduct extensive experiments showing that when consolidating with the off-policy learning framework, RPG substantially reduces the sample complexity, comparing to the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa1f10b8-e072-4e5c-903e-ac2fb9bab589Cited by top-tier papers2
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
- Learning Value Functions in Deep Policy Gradients using Residual VarianceYannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe PreuxICLR 2021 · 16 citations
Builds on1
Related papers
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
- Occupancy-based Policy Gradient: Estimation, Convergence, and OptimalityAudrey Huang, Nan JiangNeurIPS 2024 · 5 citations
- SAPG: Split and Aggregate Policy GradientsJayesh Singla, Ananye Agarwal, Deepak PathakICML 2024 · 19 citations
- Optimistic Natural Policy Gradient: a Simple Efficient Policy Optimization Framework for Online RLQinghua Liu, Gellért Weisz, András György, Chi Jin et al.NeurIPS 2023 · 16 citations
- Reinforcement Learning to Rank with Pairwise Policy GradientJun Xu, Zeng Wei, Long Xia, Yanyan Lan et al.SIGIR 2020 · 32 citations
