Multi-Agent First Order Constrained Optimization in Policy Space
Youpeng Zhao, Yaodong Yang, Zhenbo Lu, Wengang Zhou, Houqiang Li
摘要
In the realm of multi-agent reinforcement learning (MARL), achieving high performance is crucial for a successful multi-agent system. Meanwhile, the ability to avoid unsafe actions is becoming an urgent and imperative problem to solve for real-life applications. Whereas, it is still challenging to develop a safety-aware method for multi-agent systems in MARL. In this work, we introduce a novel approach called Multi-Agent First Order Constrained Optimization in Policy Space (MAFOCOPS), which effectively addresses the dual objectives of attaining satisfactory performance and enforcing safety constraints. Using data generated from the current policy, MAFOCOPS first finds the optimal update policy by solving a constrained optimization problem in the nonparameterized policy space. Then, the update policy is projected back into the parametric policy space to achieve a feasible policy. Notably, our method is first-order in nature, ensuring the ease of implementation, and exhibits an approximate upper bound on the worst-case constraint violation. Empirical results show that our approach achieves remarkable performance while satisfying safe constraints on several safe MARL benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsH. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja 等NeurIPS 2025
- Discrete GCBF Proximal Policy Optimization for Multi-agent Safe Optimal ControlSongyuan Zhang, Oswin So, Mitchell Black, Chuchu FanICLR 2025
- Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy OptimizationHao Zhang, Yaru Niu, Yikai Wang, Ding Zhao 等ICML 2026
它引用的顶会 Paper7
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 被引用 171 次
相关 Paper
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang 等NeurIPS 2022 · 被引用 95 次
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song 等NeurIPS 2024 · 被引用 22 次
- Enhancing Efficiency of Safe Reinforcement Learning via Sample ManipulationShangding Gu, Laixi Shi, Yuhao Ding, Alois Knoll 等NeurIPS 2024 · 被引用 14 次
- CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement LearningAyoub Belouadah, Sylvain Kubler, YVES LE TRAONICML 2026
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu 等NeurIPS 2025 · 被引用 1 次
