ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
Akhil Agnihotri, Rahul Jain, Haipeng Luo
Abstract
Reinforcement Learning (RL) for constrained MDPs (CMDPs) is an increasingly important problem for various applications. Often, the average criterion is more suitable than the discounted criterion. Yet, RL for average-CMDPs (ACMDPs) remains a challenging problem. Algorithms designed for discounted constrained RL problems often do not perform well for the average CMDP setting. In this paper, we introduce a new policy optimization with function approximation algorithm for constrained MDPs with the average criterion. The Average-Constrained Policy Optimization (ACPO) algorithm is inspired by trust region-based policy optimization algorithms. We develop basic sensitivity theory for average CMDPs, and then use the corresponding bounds in the design of the algorithm. We provide theoretical guarantees on its performance, and through extensive experimental work in various challenging OpenAI Gym environments, show its superior empirical performance when compared to other state-of-the-art algorithms adapted for the ACMDPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c6043c6-eb65-4d9d-b73d-08df51c17ef0Cited by top-tier papers4
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 15 citations
- e-COP : Episodic Constrained Optimization of PoliciesAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil SinglaNeurIPS 2024 · 2 citations
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
Builds on7
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm et al.ICLR 2024 · 89 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Investigating the Impact of Multi-LiDAR Placement on Object Detection for Autonomous DrivingHanjiang Hu, Zuxin Liu, Sharad Chitlangia, Akhil Agnihotri et al.CVPR 2022 · 60 citations
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
Related papers
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Constrained Offline Policy OptimizationNicholas Polosky, Bruno C. da Silva, Madalina Fiterau, Jithin JagannathICML 2022 · 21 citations
- Learning to Constrain Policy Optimization with Virtual Trust RegionHung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen et al.NeurIPS 2022 · 5 citations
