Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning
Kristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton, Daniel Graves
摘要
We explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a fixed number of future time steps. To learn the value function for horizon h, these algorithms bootstrap from the value function for horizon h−1, or some shorter horizon. Because no value function bootstraps from itself, fixed-horizon methods are immune to the stability problems that plague other off-policy TD methods using function approximation (also known as “the deadly triad”). Although fixed-horizon methods require the storage of additional value functions, this gives the agent additional predictive power, while the added complexity can be substantially reduced via parallel updates, shared weights, and n-step bootstrapping. We show how to use fixed-horizon value functions to solve reinforcement learning problems competitively with methods such as Q-learning that learn conventional value functions. We also prove convergence of fixed-horizon temporal difference methods with linear and general function approximation. Taken together, our results establish fixed-horizon TD methods as a viable new way of avoiding the stability problems of the deadly triad.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Serial Scaling HypothesisYuxi Liu, Konpat Preechakul, Kananart Kuwaranancharoen, Yutong BaiICLR 2026 · 被引用 12 次
- Chaining Value Functions for Off-Policy LearningSimon Schmitt, John Shawe-Taylor, Hado van HasseltAAAI 2022 · 被引用 5 次
相关 Paper
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 被引用 39 次
- The Pitfalls of Regularization in Off-Policy TD LearningGaurav Manek, J. Zico KolterNeurIPS 2022 · 被引用 7 次
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White 等ICML 2021 · 被引用 22 次
