Policy Gradients Incorporating the Future
David Venuto, Elaine Lau, Doina Precup, Ofir Nachum
Abstract
Reasoning about the future -- understanding how decisions in the present time affect outcomes in the future -- is one of the central challenges for reinforcement learning (RL), especially in highly-stochastic or partially observable environments. While predicting the future directly is hard, in this work we introduce a method that allows an agent to"look into the future"without explicitly predicting it. Namely, we propose to allow an agent, during its training on past experience, to observe what actually happened in the future at that time, while enforcing an information bottleneck to avoid the agent overly relying on this privileged information. This gives our agent the opportunity to utilize rich and useful information about the future trajectory dynamics in addition to the present. Our method, Policy Gradients Incorporating the Future (PGIF), is easy to implement and versatile, being applicable to virtually any policy gradient algorithm. We apply our proposed method to a number of off-the-shelf RL algorithms and show that PGIF is able to achieve higher reward faster in a variety of online and offline RL domains, as well as sparse-reward and partially observable environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea414f2e-3f25-4860-bde9-d532680d88aaCited by top-tier papers5
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Future-conditioned Unsupervised Pretraining for Decision TransformerZhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu et al.ICML 2023 · 32 citations
- Hindsight Learning for MDPs with Exogenous InputsSean R. Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall et al.ICML 2023 · 31 citations
- Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysisAlexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw et al.NeurIPS 2023 · 16 citations
- Learning to Explore in POMDPs with Informational RewardsAnnie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong et al.ICML 2024 · 2 citations
Builds on8
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
Related papers
- The Value of Reward Lookahead in Reinforcement LearningNadav Merlis, Dorian Baudry, Vianney PerchetNeurIPS 2024 · 6 citations
- Posterior Value Functions: Hindsight Baselines for Policy Gradient MethodsChris Nota, Philip S. Thomas, Bruno C. da SilvaICML 2021 · 6 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Reinforcement Learning with Lookahead InformationNadav MerlisNeurIPS 2024 · 11 citations
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without ForgettingJorge A. Mendez, Boyu Wang, Eric EatonNeurIPS 2020 · 42 citations
