Learning and Planning in Average-Reward Markov Decision Processes
Yi Wan, Abhishek Naik, Richard S. Sutton
摘要
We introduce improved learning and planning algorithms for average-reward MDPs, including 1) the first general proven-convergent off-policy model-free control algorithm without reference states, 2) the first proven-convergent off-policy model-free prediction algorithm, and 3) the first learning algorithms that converge to the actual value function rather than to the value function plus an offset. All of our algorithms are based on using the temporal-difference error rather than the conventional error when updating the estimate of the average reward. Our proof techniques are based on those of Abounadi, Bertsekas, and Borkar (2001). Empirically, we show that the use of the temporal-difference error generally results in faster learning, and that reliance on a reference state generally results in slower learning and risks divergence. All of our learning algorithms are fully online, and all of our planning algorithms are fully incremental.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- On-Policy Deep Reinforcement Learning for the Average-Reward CriterionYiming Zhang, Keith W. RossICML 2021 · 被引用 59 次
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah 等AAAI 2021 · 被引用 54 次
- Markovian Interference in ExperimentsVivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew ZhengNeurIPS 2022 · 被引用 52 次
- Finite Sample Analysis of Average-Reward TD Learning and -LearningSheng Zhang, Zhe Zhang, Siva Theja MaguluriNeurIPS 2021 · 被引用 48 次
它引用的顶会 Paper1
相关 Paper
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 被引用 12 次
- Performance Bounds for Policy-Based Average Reward Reinforcement Learning AlgorithmsYashaswini Murthy, Mehrdad Moharrami, R. SrikantNeurIPS 2023 · 被引用 10 次
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 被引用 4 次
- Off-Policy Average Reward Actor-Critic with Deterministic Policy SearchNaman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh BhatnagarICML 2023 · 被引用 14 次
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
