Learning Expected Emphatic Traces for Deep RL
Ray Jiang, Shangtong Zhang, Veronica Chelu, Adam White, Hado van Hasselt
摘要
Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approximation, such as neural networks, this combination is known as the deadly triad and is potentially unstable. Recently, it has been shown that stability and good performance at scale can be achieved by combining emphatic weightings and multi-step updates. This approach, however, is generally limited to sampling complete trajectories in order, to compute the required emphatic weighting. In this paper we investigate how to combine emphatic weightings with non-sequential, off-line data sampled from a replay buffer. We develop a multi-step emphatic weighting that can be combined with replay, and a time-reversed n-step TD learning algorithm to learn the required emphatic weighting. We show that these state weightings reduce variance compared with prior approaches, while providing convergence guarantees. We tested the approach at scale on Atari 2600 video games, and observed that the new X-ETD(n) agent improved over baseline agents, highlighting both the scalability and broad applicability of our approach. Many deep reinforcement learning systems are not sample efficient. A simple and effective way to improve sample efficiency is to make better use of prior experience via replay (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Correcting discount-factor mismatch in on-policy policy gradient methodsFengdi Che, Gautham Vasan, A. Rupam MahmoodICML 2023 · 被引用 10 次
- The Pitfalls of Regularization in Off-Policy TD LearningGaurav Manek, J. Zico KolterNeurIPS 2022 · 被引用 7 次
- PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning MethodZiwei Guan, Tengyu Xu, Yingbin LiangICLR 2022 · 被引用 5 次
- Adaptive Interest for Emphatic Reinforcement LearningMartin Klissarov, Rasool Fakoor, Jonas W. Mueller, Kavosh Asadi 等NeurIPS 2022 · 被引用 3 次
它引用的顶会 Paper10
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 125 次
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 被引用 107 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez 等ICML 2021 · 被引用 69 次
相关 Paper
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White 等ICML 2021 · 被引用 22 次
- Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningKristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton 等AAAI 2020 · 被引用 34 次
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 被引用 10 次
- Averaging n-step Returns Reduces Variance in Reinforcement LearningBrett Daley, Martha White, Marlos C. MachadoICML 2024 · 被引用 7 次
