Skill or Luck? Return Decomposition via Advantage Functions
Hsiao-Ru Pan, Bernhard Schölkopf
Abstract
Learning from off-policy data is essential for sample-efficient reinforcement learning. In the present work, we build on the insight that the advantage function can be understood as the causal effect of an action on the return, and show that this allows us to decompose the return of a trajectory into parts caused by the agent's actions (skill) and parts outside of the agent's control (luck). Furthermore, this decomposition enables us to naturally extend Direct Advantage Estimation (DAE) to off-policy settings (Off-policy DAE). The resulting method can learn from off-policy trajectories without relying on importance sampling techniques or truncating off-policy actions. We draw connections between Off-policy DAE and previous methods to demonstrate how it can speed up learning and when the proposed off-policy corrections are important. Finally, we use the MinAtar environments to illustrate how ignoring off-policy corrections can lead to suboptimal policy optimization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd51ac6f-bdcd-415a-ad2a-4f41a77e6a31Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert et al.ICLR 2022 · 79 citations
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
Related papers
- Direct Advantage EstimationHsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard SchölkopfNeurIPS 2022 · 20 citations
- VA-learning as a more efficient alternative to Q-learningYunhao Tang, Rémi Munos, Mark Rowland, Michal ValkoICML 2023 · 11 citations
- Truncating Trajectories in Monte Carlo Reinforcement LearningRiccardo Poiani, Alberto Maria Metelli, Marcello RestelliICML 2023 · 6 citations
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal et al.ICLR 2021 · 3 citations
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
