Lune

NeurIPS2024Top-tier venue

e-COP : Episodic Constrained Optimization of Policies

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil Singla

2024Year
2Citations

Abstract

In this paper, we present the e-COP\texttt{e-COP} algorithm, the first policy optimization algorithm for constrained Reinforcement Learning (RL) in episodic (finite horizon) settings. Such formulations are applicable when there are separate sets of optimization criteria and constraints on a system's behavior. We approach this problem by first establishing a policy difference lemma for the episodic setting, which provides the theoretical foundation for the algorithm. Then, we propose to combine a set of established and novel solution ideas to yield the e-COP\texttt{e-COP} algorithm that is easy to implement and numerically stable, and provide a theoretical guarantee on optimality under certain scaling assumptions. Through extensive empirical analysis using benchmarks in the Safety Gym suite, we show that our algorithm has similar or better performance than SoTA (non-episodic) algorithms adapted for the episodic setting. The scalability of the algorithm opens the door to its application in safety-constrained Reinforcement Learning from Human Feedback for Large Language or Diffusion Models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext eb5f8687-b0f2-48da-93ea-6e6eff0f4aef

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines