Discerning Temporal Difference Learning
Jianfei Ma
摘要
Temporal difference learning (TD) is a foundational concept in reinforcement learning (RL), aimed at efficiently assessing a policy's value function. TD(λ), a potent variant, incorporates a memory trace to distribute the prediction error into the historical context. However, this approach often neglects the significance of historical states and the relative importance of propagating the TD error, influenced by challenges such as visitation imbalance or outcome noise. To address this, we propose a novel TD algorithm named discerning TD learning (DTD), which allows flexible emphasis functions—predetermined or adapted during training—to allocate efforts effectively across states. We establish the convergence properties of our method within a specific class of emphasis functions and showcase its promising potential for adaptation to deep RL contexts. Empirical results underscore that employing a judicious emphasis function not only improves value estimation but also expedites learning across diverse scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 被引用 85 次
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon 等AAAI 2020 · 被引用 51 次
- Selective Dyna-Style Planning Under Limited Model CapacityZaheer Abbas, Samuel Sokota, Erin Talvitie, Martha WhiteICML 2020 · 被引用 38 次
- Continual Auxiliary Task LearningMatthew McLeod, Chunlok Lo, Matthew Schlegel, Andrew Jacobsen 等NeurIPS 2021 · 被引用 13 次
相关 Paper
- Preferential Temporal Difference LearningNishanth V. Anand, Doina PrecupICML 2021 · 被引用 9 次
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White 等ICML 2021 · 被引用 22 次
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman 等NeurIPS 2023 · 被引用 17 次
- Mitigating Partial Observability in Sequential Decision Processes via the Lambda DiscrepancyCameron Allen, Aaron Kirtland, Ruo Yu Tao, Sam Lobel 等NeurIPS 2024 · 被引用 12 次
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska 等ICML 2022 · 被引用 40 次
