Quantile Credit Assignment
Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, Rémi Munos
Abstract
In reinforcement learning, the credit assignment problem is to distinguish luck from skill, that is, separate the inherent randomness in the environment from the controllable effects of the agent's actions. This paper proposes two novel algorithms, Quantile Credit Assignment (QCA) and Hindsight QCA (HQCA), which incorporate distributional value estimation to perform credit assignment. QCA uses a network that predicts the quantiles of the return distribution, whereas HQCA additionally incorporates information about the future. Both QCA and HQCA have the appealing interpretation of leveraging an estimate of the quantile level of the return (interpreted as the level of "luck") in order to derive a "luck-dependent" baseline for policy gradient methods. We show theoretically that this approach gives an unbiased policy gradient estimator that can yield significant variance reductions over a standard value estimate baseline. QCA and HQCA significantly outperform prior state-ofthe-art methods on a range of extremely difficult credit assignment problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3ce889d-df42-4b46-9438-bd749bdce53cCited by top-tier papers1
Ask how each one uses itBuilds on4
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Non-Crossing Quantile Regression for Distributional Reinforcement LearningFan Zhou, Jianing Wang, Xingdong FengNeurIPS 2020 · 63 citations
- GMAC: A Distributional Perspective on Actor-Critic FrameworkDaniel Wontae Nam, Younghoon Kim, Chan Y. ParkICML 2021 · 22 citations
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte et al.ICLR 2022 · 18 citations
Related papers
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 46 citations
- Hindsight Network Credit Assignment: Efficient Credit Assignment in Networks of Discrete Stochastic UnitsKenny YoungAAAI 2022
- Posterior Value Functions: Hindsight Baselines for Policy Gradient MethodsChris Nota, Philip S. Thomas, Bruno C. da SilvaICML 2021 · 6 citations
- Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysisAlexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw et al.NeurIPS 2023 · 16 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
