Off-Policy Average Reward Actor-Critic with Deterministic Policy Search
Naman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh Bhatnagar
Abstract
The average reward criterion is relatively less studied as most existing works in the Reinforcement Learning literature consider the discounted reward criterion. There are few recent works that present on-policy average reward actor-critic algorithms, but average reward off-policy actor-critic is relatively less explored. In this work, we present both on-policy and off-policy deterministic policy gradient theorems for the average reward performance criterion. Using these theorems, we also present an Average Reward Off-Policy Deep Deterministic Policy Gradient (ARO-DDPG) Algorithm. We first show asymptotic convergence analysis using the ODE-based method. Subsequently, we provide a finite time analysis of the resulting stochastic approximation scheme with linear function approximator and obtain an -optimal stationary policy with a sample complexity of . We compare the average reward performance of our proposed ARO-DDPG algorithm and observe better empirical performance compared to state-of-the-art on-policy average reward actor-critic algorithms over MuJoCo-based environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba5a5532-b9de-4d44-b4cd-70dfebaeacb5Cited by top-tier papers5
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao et al.ICML 2024 · 15 citations
- NeoRL: Efficient Exploration for Nonepisodic RLBhavya Sukhija, Lenart Treven, Florian Dörfler, Stelian Coros et al.NeurIPS 2024 · 7 citations
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 6 citations
- Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian NoiseEthan Blaser, Shangtong ZhangAAAI 2026 · 3 citations
- Efficient Offline Reinforcement Learning via Peer-Influenced ConstraintYujia Zhang, Lin Li, Wei Wei, Jianguo Wu et al.ICLR 2026
Builds on8
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 189 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 107 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
Related papers
- On-Policy Deep Reinforcement Learning for the Average-Reward CriterionYiming Zhang, Keith W. RossICML 2021 · 59 citations
- A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic ApproachSwetha Ganesh, Washim Uddin Mondal, Vaneet AggarwalICML 2025
- Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network ApproximationXuyang Chen, Fengzhuo Zhang, Keyu Yan, Lin ZhaoICLR 2026
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 1 citation
- Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time OraclesBhrij Patel, Wesley A. Suttle, Alec Koppel, Vaneet Aggarwal et al.ICML 2024 · 4 citations
