Lune

NeurIPS2024Top-tier venue

Sample-Efficient Constrained Reinforcement Learning with General Parameterization

Washim Uddin Mondal, Vaneet Aggarwal

2024Year
15Citations
4Top-tier citations

Abstract

We consider a constrained Markov Decision Problem (CMDP) where the goal of an agent is to maximize the expected discounted sum of rewards over an infinite horizon while ensuring that the expected discounted sum of costs exceeds a certain threshold. Building on the idea of momentum-based acceleration, we develop the Primal-Dual Accelerated Natural Policy Gradient (PD-ANPG) algorithm that ensures an ϵ\epsilon global optimality gap and ϵ\epsilon constraint violation with O~((1−γ)−7ϵ−2)\tilde{\mathcal{O}}((1-\gamma)^{-7}\epsilon^{-2}) sample complexity for general parameterized policies where γ\gamma denotes the discount factor. This improves the state-of-the-art sample complexity in general parameterized CMDPs by a factor of O((1−γ)−1ϵ−2)\mathcal{O}((1-\gamma)^{-1}\epsilon^{-2}) and achieves the theoretical lower bound in ϵ−1\epsilon^{-1}.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d0914bf2-3098-41a2-b99c-b3e8ee97b0e3

Cited by top-tier papers4

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines