Shield Decentralization for Safe Multi-Agent Reinforcement Learning
Daniel Melcer, Christopher Amato, Stavros Tripakis
Abstract
Learning safe solutions is an important but challenging problem in multi-agent re-inforcement learning (MARL). Shielded reinforcement learning is one approach for preventing agents from choosing unsafe actions. Current shielded reinforcement learning methods for MARL make strong assumptions about communication and full observability. In this work, we extend the formalization of the shielded reinforcement learning problem to a decentralized multi-agent setting. We then present an algorithm for decomposition of a centralized shield, allowing shields to be used in such decentralized, communication-free environments. Our results show that agents equipped with decentralized shields perform comparably to agents with centralized shields in several tasks, allowing shielding to be used in environments with decentralized training and execution for the first time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0b3ac32-a047-4baf-8d33-234a8a529badCited by top-tier papers8
- Safe Exploration in Reinforcement Learning: A Generalized Formulation and AlgorithmsAkifumi Wachi, Wataru Hashimoto, Xun Shen, Kazumune HashimotoNeurIPS 2023 · 38 citations
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesRunze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros et al.NeurIPS 2025 · 7 citations
- HypRL: Reinforcement Learning of Control Policies for HyperpropertiesTzu-Han Hsu, Arshia Rafieioskouei, Borzoo BonakdarpourNeurIPS 2025 · 5 citations
- Flipping-based Policy for Chance-Constrained Markov Decision ProcessesXun Shen, Shuo Jiang, Akifumi Wachi, Kazumune Hashimoto et al.NeurIPS 2024 · 5 citations
Related papers
- Consensus Learning for Cooperative Multi-Agent Reinforcement LearningZhiwei Xu, Bin Zhang, Dapeng Li, Zeren Zhang et al.AAAI 2023 · 27 citations
- MA2E: Addressing Partial Observability in Multi-Agent Reinforcement Learning with Masked Auto-EncoderSehyeok Kang, Yongsik Lee, Gahee Kim, Song Chong et al.ICLR 2025
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 7 citations
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar et al.AAAI 2021 · 93 citations
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar et al.NeurIPS 2021 · 105 citations
