Learning and Planning in Average-Reward Markov Decision Processes
Yi Wan, Abhishek Naik, Richard S. Sutton
Abstract
We introduce improved learning and planning algorithms for average-reward MDPs, including 1) the first general proven-convergent off-policy model-free control algorithm without reference states, 2) the first proven-convergent off-policy model-free prediction algorithm, and 3) the first learning algorithms that converge to the actual value function rather than to the value function plus an offset. All of our algorithms are based on using the temporal-difference error rather than the conventional error when updating the estimate of the average reward. Our proof techniques are based on those of Abounadi, Bertsekas, and Borkar (2001). Empirically, we show that the use of the temporal-difference error generally results in faster learning, and that reliance on a reference state generally results in slower learning and risks divergence. All of our learning algorithms are fully online, and all of our planning algorithms are fully incremental.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb66f071-fb91-4fb5-846a-6f4329fd6a3eCited by top-tier papers26
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- On-Policy Deep Reinforcement Learning for the Average-Reward CriterionYiming Zhang, Keith W. RossICML 2021 · 59 citations
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah et al.AAAI 2021 · 54 citations
- Markovian Interference in ExperimentsVivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew ZhengNeurIPS 2022 · 52 citations
- Finite Sample Analysis of Average-Reward TD Learning and -LearningSheng Zhang, Zhe Zhang, Siva Theja MaguluriNeurIPS 2021 · 48 citations
Builds on1
Related papers
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 12 citations
- Performance Bounds for Policy-Based Average Reward Reinforcement Learning AlgorithmsYashaswini Murthy, Mehrdad Moharrami, R. SrikantNeurIPS 2023 · 10 citations
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Off-Policy Average Reward Actor-Critic with Deterministic Policy SearchNaman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh BhatnagarICML 2023 · 14 citations
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
