On-Policy Deep Reinforcement Learning for the Average-Reward Criterion
Yiming Zhang, Keith W. Ross
Abstract
We develop theory and algorithms for average-reward on-policy Reinforcement Learning (RL). We first consider bounding the difference of the long-term average reward for two policies. We show that previous work based on the discounted return (Schulman et al., 2015; Achiam et al., 2017) results in a non-meaningful bound in the average-reward setting. By addressing the average-reward criterion directly, we then derive a novel bound which depends on the average divergence between the two policies and Kemeny's constant. Based on this bound, we develop an iterative procedure which produces a sequence of monotonically improved policies for the average reward criterion. This iterative procedure can then be combined with classic DRL (Deep Reinforcement Learning) methods, resulting in practical DRL algorithms that target the long-run average reward criterion. In particular, we demonstrate that Average-Reward TRPO (ATRPO), which adapts the on-policy TRPO algorithm to the average-reward criterion, significantly outperforms TRPO in the most challenging MuJuCo environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97467f74-3abc-4f97-86a3-849730c294bcCited by top-tier papers12
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 47 citations
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 25 citations
- Off-Policy Average Reward Actor-Critic with Deterministic Policy SearchNaman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh BhatnagarICML 2023 · 14 citations
- NeoRL: Efficient Exploration for Nonepisodic RLBhavya Sukhija, Lenart Treven, Florian Dörfler, Stelian Coros et al.NeurIPS 2024 · 7 citations
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 6 citations
Builds on5
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision ProcessesChen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma et al.ICML 2020 · 120 citations
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 85 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
Related papers
- Performance Bounds for Policy-Based Average Reward Reinforcement Learning AlgorithmsYashaswini Murthy, Mehrdad Moharrami, R. SrikantNeurIPS 2023 · 10 citations
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 16 citations
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou et al.ICML 2021 · 24 citations
