Near-Optimal Multi-Agent Learning for Safe Coverage Control
Manish Prajapat, Matteo Turchetta, Melanie N. Zeilinger, Andreas Krause
Abstract
In multi-agent coverage control problems, agents navigate their environment to reach locations that maximize the coverage of some density. In practice, the density is rarely known a priori, further complicating the original NP-hard problem. Moreover, in many applications, agents cannot visit arbitrary locations due to a priori unknown safety constraints. In this paper, we aim to efficiently learn the density to approximately solve the coverage problem while preserving the agents' safety. We first propose a conditionally linear submodular coverage function that facilitates theoretical analysis. Utilizing this structure, we develop MACOPT, a novel algorithm that efficiently trades off the exploration-exploitation dilemma due to partial observability, and show that it achieves sublinear regret. Next, we extend results on single-agent safe exploration to our multi-agent setting and propose SAFEMAC for safe coverage and exploration. We analyze SAFEMAC and give first of its kind results: near optimal coverage in finite time while provably guaranteeing safety. We extensively evaluate our algorithms on synthetic and real problems, including a biodiversity monitoring task under safety constraints, where SAFEMAC outperforms competing methods. Introduction In multi-agent coverage control (MAC) problems, multiple agents coordinate to maximize coverage over some spatially distributed events. Their applications abound, from collaborative mapping [1], environmental monitoring [2], inspection robotics [3] to sensor networks [4] . In addition, the coverage formulation can address core challenges in cooperative multi-agent RL [5, 6] , e.g., exploration [7], by providing high-level goals. In these applications, agents often encounter safety constraints that may lead to critical accidents when ignored, e.g., obstacles [8] or extreme weather conditions [9, 10] . Deploying coverage control solutions in the real world presents many challenges: (i) for a given density of relevant events, this is an NP hard problem [11] ; (ii) such density is rarely known in practice [2] and must be learned from data, which presents a complex active learning problem as the quantity we measure (the density) differs from the one we want to optimize (its coverage); (iii) agents often operate under safety-critical conditions, [8-10], that may be unknown a priori. This requires cautious exploration of the environment to prevent catastrophic outcomes. While prior work addresses subsets of these challenges (see Section 7), we are not aware of methods that address them jointly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Submodular Reinforcement LearningManish Prajapat, Mojmir Mutny, Melanie N. Zeilinger, Andreas KrauseICLR 2024 · 26 citations
- Global Reinforcement Learning : Beyond Linear and Convex Rewards via Submodular Semi-gradient MethodsRiccardo De Santi, Manish Prajapat, Andreas KrauseICML 2024 · 14 citations
- Finding Safe Zones of Markov Decision Processes PoliciesLee Cohen, Yishay Mansour, Michal MoshkovitzNeurIPS 2023 · 1 citation
- Learning Safe Control via On-the-Fly Bandit ExplorationAlexandre Capone, Ryan Kazuo Cosner, Aaron D. Ames, Sandra HircheICML 2025
- Near-Optimal Online Learning for Multi-Agent Submodular Coordination: Tight Approximation and Communication EfficiencyQixin Zhang, Zongqi Wan, Yu Yang, Li Shen et al.ICLR 2025
Builds on5
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause et al.NeurIPS 2020 · 109 citations
- Experimental Design for Linear Functionals in Reproducing Kernel Hilbert SpacesMojmir Mutny, Andreas KrauseNeurIPS 2022 · 12 citations
Related papers
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 24 citations
- Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor NetworksJing Xu, Fangwei Zhong, Yizhou WangNeurIPS 2020 · 71 citations
- Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement LearningOswin So, Eric Yang Yu, Songyuan Zhang, Matthew Cleaveland et al.ICLR 2026 · 1 citation
- Multi-Agent Reinforcement Learning with Submodular RewardWenjing Chen, Chengyuan Qian, Shuo Xing, Yi Zhou et al.ICML 2026 · 2 citations
