Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement Learning
Songtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar, Lior Horesh
Abstract
This paper deals with distributed reinforcement learning problems with safety constraints. In particular, we consider that a team of agents cooperate in a shared environment, where each agent has its individual reward function and safety constraints that involve all agents' joint actions. As such, the agents aim to maximize the team-average long-term return, subject to all the safety constraints. More intriguingly, no central controller is assumed to coordinate the agents, and both the rewards and constraints are only known to each agent locally/privately. Instead, the agents are connected by a peer-to-peer communication network to share information with their neighbors. In this work, we first formulate this problem as a distributed constrained Markov decision process (D-CMDP) with networked agents. Then, we propose a decentralized policy gradient (PG) method, Safe Dec-PG, to perform policy optimization based on this D-CMDP model over a network. Convergence guarantees, together with numerical results, showcase the superiority of the proposed algorithm. To the best of our knowledge, this is the first decentralized PG algorithm that accounts for the coupled safety constraints with a quantifiable convergence rate in multi-agent reinforcement learning. Finally, we emphasize that our algorithm is also novel in solving a class of decentralized stochastic nonconvex-concave minimax optimization problems, where both the algorithm design and corresponding theoretical analysis are of independent interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aeb55c02-83c1-41b3-a768-ecc74ff5d77eCited by top-tier papers16
- A Faster Decentralized Algorithm for Nonconvex Minimax ProblemsWenhan Xian, Feihu Huang, Yanfu Zhang, Heng HuangNeurIPS 2021 · 72 citations
- Taming Communication and Sample Complexities in Decentralized Policy Evaluation for Cooperative Multi-Agent Reinforcement LearningXin Zhang, Zhuqing Liu, Jia Liu, Zhengyuan Zhu et al.NeurIPS 2021 · 36 citations
- Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General UtilitiesDonghao Ying, Yunkai Zhang, Yuhao Ding, Alec Koppel et al.NeurIPS 2023 · 28 citations
- PrefPaint: Aligning Image Inpainting Diffusion Model with Human PreferenceKendong Liu, Zhiyu Zhu, Chuanhao Li, Hui Liu et al.NeurIPS 2024 · 26 citations
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
Builds on4
- On Gradient Descent Ascent for Nonconvex-Concave Minimax ProblemsTianyi Lin, Chi Jin, Michael I. JordanICML 2020 · 587 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- A Decentralized Parallel Algorithm for Training Generative Adversarial NetsMingrui Liu, Wei Zhang, Youssef Mroueh, Xiaodong Cui et al.NeurIPS 2020 · 6 citations
Related papers
- DeCOM: Decomposed Policy for Constrained Cooperative Multi-Agent Reinforcement LearningZhaoxing Yang, Haiming Jin, Rong Ding, Haoyi You et al.AAAI 2023 · 5 citations
- Learning Nash Equilibrium of Markov Potential Games with a Shared Constraint via Primal-Dual OptimizationSongtao Feng, Michael R. Dorothy, Jie FuAAAI 2025
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement LearningDohyeong Kim, Mineui Hong, Jeongho Park, Songhwai OhICLR 2025
- Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual AlgorithmQinbo Bai, Amrit Singh Bedi, Vaneet AggarwalAAAI 2023 · 29 citations
