ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
Akhil Agnihotri, Rahul Jain, Haipeng Luo
摘要
Reinforcement Learning (RL) for constrained MDPs (CMDPs) is an increasingly important problem for various applications. Often, the average criterion is more suitable than the discounted criterion. Yet, RL for average-CMDPs (ACMDPs) remains a challenging problem. Algorithms designed for discounted constrained RL problems often do not perform well for the average CMDP setting. In this paper, we introduce a new policy optimization with function approximation algorithm for constrained MDPs with the average criterion. The Average-Constrained Policy Optimization (ACPO) algorithm is inspired by trust region-based policy optimization algorithms. We develop basic sensitivity theory for average CMDPs, and then use the corresponding bounds in the design of the algorithm. We provide theoretical guarantees on its performance, and through extensive experimental work in various challenging OpenAI Gym environments, show its superior empirical performance when compared to other state-of-the-art algorithms adapted for the ACMDPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 被引用 15 次
- e-COP : Episodic Constrained Optimization of PoliciesAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil SinglaNeurIPS 2024 · 被引用 2 次
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
它引用的顶会 Paper7
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm 等ICLR 2024 · 被引用 89 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
- Investigating the Impact of Multi-LiDAR Placement on Object Detection for Autonomous DrivingHanjiang Hu, Zuxin Liu, Sharad Chitlangia, Akhil Agnihotri 等CVPR 2022 · 被引用 60 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
相关 Paper
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Constrained Offline Policy OptimizationNicholas Polosky, Bruno C. da Silva, Madalina Fiterau, Jithin JagannathICML 2022 · 被引用 21 次
- Learning to Constrain Policy Optimization with Virtual Trust RegionHung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen 等NeurIPS 2022 · 被引用 5 次
