Lune

NeurIPS2025Top-tier venue

Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm

Yang Xu, Swetha Ganesh, Washim Uddin Mondal, Qinbo Bai, Vaneet Aggarwal

2025Year
8Citations
1Top-tier citations

Abstract

This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) with general parametrization. We propose a Primal-Dual Natural Actor-Critic algorithm that adeptly manages constraints while ensuring a high convergence rate. In particular, our algorithm achieves global convergence and constraint violation rates of O~(1/T)\tilde{\mathcal{O}}(1/\sqrt{T}) over a horizon of length TT when the mixing time, τmix\tau_{\mathrm{mix}}, is known to the learner. In absence of knowledge of τmix\tau_{\mathrm{mix}}, the achievable rates change to O~(1/T0.5−ϵ)\tilde{\mathcal{O}}(1/T^{0.5-\epsilon}) provided that T≥O~(τmix2/ϵ)T \geq \tilde{\mathcal{O}}\left(\tau_{\mathrm{mix}}^{2/\epsilon}\right). Our results match the theoretical lower bound for Markov Decision Processes and establish a new benchmark in the theoretical exploration of average reward CMDPs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 104dabf9-90b0-4ca7-be7c-ebf546843e16

Cited by top-tier papers1

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines