Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement Learning
Lijun Zhang, Lin Li, Wei Wei, Huizhong Song, Yaodong Yang, Jiye Liang
摘要
. Abstract A challenging problem in seeking to bring multi-agent reinforcement learning (MARL) techniques into real-world applications, such as autonomous driving and drone swarms, is how to control multiple agents safely and cooperatively to accomplish tasks. Most existing safe MARL methods learn the centralized value function by introducing a global state to guide safety cooperation. However, the global coupling arising from safety constraints and the exponential growth of the state-action space size limit their applicability in instant communication or computing resource-constrained systems and larger multi-agent systems. In this paper, we develop a novel scalable and theoretically-justified multi-agent constrained policy optimization method. This method integrates the rigorous bounds of the trust region method and the bounds of the truncated advantage function to provide a new local policy optimization objective for each agent. Also, we prove that the safety constraints and the joint policy improvement can be met when each agent adopts a sequential update scheme to optimize a κ -hop policy. Furthermore, we propose a practical algorithm called Scalable MAPPO-Lagrangian (Scal-MAPPO-L). The proposed method’s effectiveness is verified on a collection of benchmark tasks, and the results support our theory that decentralized training with local interactions can still improve reward performance and satisfy safe constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Adversarial Attack on Black-Box Multi-Agent by Adaptive PerturbationJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie 等AAAI 2026 · 被引用 1 次
- Efficient Offline Reinforcement Learning via Peer-Influenced ConstraintYujia Zhang, Lin Li, Wei Wei, Jianguo Wu 等ICLR 2026
- Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy OptimizationQiwei Liu, Ye Yuan, Lingyue Zhang, Kaitian Chen 等AAAI 2026
- Beyond Rule-Based Agents: Active Markov Games for Realistic Multi-Agent Interaction in Autonomous DrivingYuan Gui, Hongchen Luo, Jiao Wang, Qu LiqiCVPR 2026
- HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsH. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja 等NeurIPS 2025
它引用的顶会 Paper10
- Graph Convolutional Reinforcement LearningJiechuan Jiang, Chen Dun, Tiejun Huang, Zongqing LuICLR 2020 · 被引用 415 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Multi-agent Reinforcement Learning for Networked System ControlTianshu Chu, Sandeep Chinchali, Sachin KattiICLR 2020 · 被引用 134 次
相关 Paper
- Multi-Agent First Order Constrained Optimization in Policy SpaceYoupeng Zhao, Yaodong Yang, Zhenbo Lu, Wengang Zhou 等NeurIPS 2023 · 被引用 12 次
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement LearningAyoub Belouadah, Sylvain Kubler, YVES LE TRAONICML 2026
- Augmented Proximal Policy Optimization for Safe Reinforcement LearningJuntao Dai, Jiaming Ji, Long Yang, Qian Zheng 等AAAI 2023 · 被引用 32 次
- Proactive Constrained Policy Optimization with Preemptive PenaltyNing Yang, Pengyu Wang, Guoqing Liu, Haifeng Zhang 等AAAI 2026 · 被引用 1 次
