Forethought and Hindsight in Credit Assignment
Veronica Chelu, Doina Precup, Hado van Hasselt
摘要
We address the problem of credit assignment in reinforcement learning and explore fundamental questions regarding the way in which an agent can best use additional computation to propagate new information, by planning with internal models of the world to improve its predictions. Particularly, we work to understand the gains and peculiarities of planning employed as forethought via forward models or as hindsight operating with backward models. We establish the relative merits, limitations and complementary properties of both planning mechanisms in carefully constructed scenarios. Further, we investigate the best use of models in planning, primarily focusing on the selection of states in which predictions should be (re)-evaluated. Lastly, we discuss the issue of model estimation and highlight a spectrum of methods that stretch from explicit environment-dynamics predictors to more abstract planner-aware models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Offline Reinforcement Learning with Reverse Model-based ImaginationJianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu 等NeurIPS 2021 · 被引用 74 次
- Learning Retrospective Knowledge with Reverse Reinforcement LearningShangtong Zhang, Vivek Veeriah, Shimon WhitesonNeurIPS 2020 · 被引用 13 次
- Self-Consistent Models and ValuesGregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos 等NeurIPS 2021 · 被引用 10 次
- Preferential Temporal Difference LearningNishanth V. Anand, Doina PrecupICML 2021 · 被引用 9 次
- From Past to Future: Rethinking Eligibility TracesDhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu 等AAAI 2024 · 被引用 5 次
它引用的顶会 Paper8
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function ApproximationShangtong Zhang, Bo Liu, Hengshuai Yao, Shimon WhitesonICML 2020 · 被引用 58 次
相关 Paper
- Hindsight PRIORs for Reward Learning from Human PreferencesMudit Verma, Katherine MetcalfICLR 2024 · 被引用 11 次
- Predicting Future Actions of Reinforcement Learning AgentsStephen Chung, Scott Niekum, David KruegerNeurIPS 2024 · 被引用 6 次
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang 等ICML 2023 · 被引用 3 次
- Sequence Compression Speeds Up Credit Assignment in Reinforcement LearningAditya A. Ramesh, Kenny John Young, Louis Kirsch, Jürgen SchmidhuberICML 2024 · 被引用 2 次
