From Past to Future: Rethinking Eligibility Traces
Dhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu, Philip S. Thomas, Bruno Castro da Silva
摘要
In this paper, we introduce a fresh perspective on the challenges of credit assignment and policy evaluation. First, we delve into the nuances of eligibility traces and explore instances where their updates may result in unexpected credit assignment to preceding states. From this investigation emerges the concept of a novel value function, which we refer to as the ????????????? ????? ????????. Unlike traditional state value functions, bidirectional value functions account for both future expected returns (rewards anticipated from the current state onward) and past expected returns (cumulative rewards from the episode's start to the present). We derive principled update equations to learn this value function and, through experimentation, demonstrate its efficacy in enhancing the process of policy evaluation. In particular, our results indicate that the proposed learning approach can, in certain challenging contexts, perform policy evaluation more rapidly than TD(λ)–a method that learns forward value functions, v^π, ????????. Overall, our findings present a new perspective on eligibility traces and potential advantages associated with the novel value function it inspires, especially for policy evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Offline Reinforcement Learning with Reverse Model-based ImaginationJianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu 等NeurIPS 2021 · 被引用 74 次
- Forethought and Hindsight in Credit AssignmentVeronica Chelu, Doina Precup, Hado van HasseltNeurIPS 2020 · 被引用 29 次
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 被引用 21 次
- Learning Retrospective Knowledge with Reverse Reinforcement LearningShangtong Zhang, Vivek Veeriah, Shimon WhitesonNeurIPS 2020 · 被引用 13 次
相关 Paper
- Expected Eligibility TracesHado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver 等AAAI 2021 · 被引用 49 次
- A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature PredictionsAnthony GX-Chen, Veronica Chelu, Blake A. Richards, Joelle PineauAAAI 2022 · 被引用 1 次
- Discerning Temporal Difference LearningJianfei MaAAAI 2024 · 被引用 2 次
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang 等ICML 2023 · 被引用 3 次
