Discrete GCBF Proximal Policy Optimization for Multi-agent Safe Optimal Control
Songyuan Zhang, Oswin So, Mitchell Black, Chuchu Fan
Abstract
Control policies that can achieve high task performance and satisfy safety constraints are desirable for any system, including multi-agent systems (MAS). One promising technique for ensuring the safety of MAS is distributed control barrier functions (CBF). However, it is difficult to design distributed CBF-based policies for MAS that can tackle unknown discrete-time dynamics, partial observability, changing neighborhoods, and input constraints, especially when a distributed high-performance nominal policy that can achieve the task is unavailable. To tackle these challenges, we propose DGPPO, a new framework that simultaneously learns both a discrete graph CBF which handles neighborhood changes and input constraints, and a distributed high-performance safe policy for MAS with unknown discrete-time dynamics. We empirically validate our claims on a suite of multi-agent tasks spanning three different simulation engines. The results suggest that, compared with existing methods, our DGPPO framework obtains policies that achieve high task performance (matching baselines that ignore the safety constraints), and high safety rates (matching the most conservative baselines), with a constant set of hyperparameters across all environments. 1 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9842824e-22e9-4f56-bd66-a2c7dd85c35dCited by top-tier papers3
- Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph FormXuefeng Wang, Lei Zhang, Henglin Pu, Husheng Li et al.ICLR 2026 · 1 citation
- HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsH. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja et al.NeurIPS 2025
- How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?Xiaoyuan Cheng, Wenxuan Yuan, Boyang Li, Yuanchao Xu et al.ICML 2026
Builds on9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Learning Safe Multi-agent Control with Decentralized Neural Barrier CertificatesZengyi Qin, Kaiqing Zhang, Yuxiao Chen, Jingkai Chen et al.ICLR 2021 · 164 citations
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar et al.AAAI 2021 · 93 citations
Related papers
- Learning Safe Control via On-the-Fly Bandit ExplorationAlexandre Capone, Ryan Kazuo Cosner, Aaron D. Ames, Sandra HircheICML 2025
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang et al.ICML 2023 · 81 citations
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- Safe Pontryagin Differentiable ProgrammingWanxin Jin, Shaoshuai Mou, George J. PappasNeurIPS 2021 · 65 citations
- Constrained Diffusers for Safe Planning and ControlJichen Zhang, Liqun Zhao, Antonis Papachristodoulou, Jack UmenbergerNeurIPS 2025 · 17 citations
