Emphatic Algorithms for Deep Reinforcement Learning
Ray Jiang, Tom Zahavy, Zhongwen Xu, Adam White, Matteo Hessel, Charles Blundell, Hado van Hasselt
Abstract
Off-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms can become unstable when combined with function approximation and off-policy sampling-this is known as the "deadly triad". Emphatic temporal difference (ETD(λ)) algorithm ensures convergence in the linear case by appropriately weighting the TD(λ) updates. In this paper, we extend the use of emphatic methods to deep reinforcement learning agents. We show that naively adapting ETD(λ) to popular deep reinforcement learning algorithms, which use forward view multi-step returns, results in poor performance. We then derive new emphatic algorithms for use in the context of such algorithms, and we demonstrate that they provide noticeable benefits in small problems designed to highlight the instability of TD methods. Finally, we observed improved performance when applying these algorithms at scale on classic Atari games from the Arcade Learning Environment. Off-policy learning, whereby an agent learns from behavior that differs from its current policy, affords an agent opportunities to accumulate rich knowledge (Degris & Modayil, 2012 ) by learning about the effect of different policies of behaviors. This can also be extended to learn about different goals, e.g., by learning general value functions (Sutton et al., 2011) for cumulants that differ from the main task reward. Unfortunately, it is well known that reinforcement learning algorithms (Sutton & Barto, 2018) can become unstable when combining function approximation, off-policy learning, and bootstrapping (Tsitsiklis & Van Roy, 1997)-for this reason such combination is referred to as the deadly triad (Sutton & Barto, 2018; van Hasselt et al., 2018) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85fd6488-ce34-409c-93db-0111414f87acCited by top-tier papers4
- Multi-Stage Episodic Control for Strategic Exploration in Text GamesJens Tuyls, Shunyu Yao, Sham M. Kakade, Karthik NarasimhanICLR 2022 · 30 citations
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White et al.AAAI 2022 · 14 citations
- Improving Offline RL by Blending HeuristicsSinong Geng, Aldo Pacchiano, Andrey Kolobov, Ching-An ChengICLR 2024 · 12 citations
- Adaptive Interest for Emphatic Reinforcement LearningMartin Klissarov, Rasool Fakoor, Jonas W. Mueller, Kavosh Asadi et al.NeurIPS 2022 · 3 citations
Builds on2
Related papers
- The Pitfalls of Regularization in Off-Policy TD LearningGaurav Manek, J. Zico KolterNeurIPS 2022 · 7 citations
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 39 citations
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 10 citations
- Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningKristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton et al.AAAI 2020 · 34 citations
