Multi Agent Reinforcement Learning for Sequential Satellite Assignment Problems
Joshua Holder, Natasha Jaques, Mehran Mesbahi
摘要
Assignment problems are a classic combinatorial optimization problem in which a group of agents must be assigned to a group of tasks such that maximum utility is achieved while satisfying assignment constraints. Given the utility of each agent completing each task, polynomial-time algorithms exist to solve a single assignment problem in its simplest form. However, in many modern-day applications such as satellite constellations, power grids, and mobile robot scheduling, assignment problems unfold over time, with the utility for a given assignment depending heavily on the state of the system. We apply multi-agent reinforcement learning to this problem, learning the value of assignments by bootstrapping from the known polynomial-time greedy solver and then learning from further experience. We then choose assignments using a distributed optimal assignment mechanism rather than by selecting them directly. We demonstrate that this algorithm is theoretically justified and avoids pitfalls experienced by other RL algorithms in this setting. Finally, we show that our algorithm significantly outperforms other methods in the literature, even while scaling to realistic scenarios with hundreds of agents and tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and MethodologyLuting Wang, Yinghao Xiang, Hongliang Huang, Dongjun Li 等NeurIPS 2025 · 被引用 6 次
- Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?Zijian Zhao, Sen LiICLR 2026 · 被引用 4 次
- OrbitZoo: Real Orbital Systems Challenges for Reinforcement LearningAlexandre Oliveira, Katarina Dyreby, Francisco M. Caldas, Cláudia SoaresNeurIPS 2025 · 被引用 2 次
- Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARLJuan Claude Formanek, Omayma Mahjoub, Louay Ben Nessir, Sasha Abramowitz 等NeurIPS 2025
它引用的顶会 Paper1
相关 Paper
- Heterogeneous Graph Transformers for Simultaneous Mobile Multi-Robot Task Allocation and Scheduling under Temporal ConstraintsBatuhan Altundas, Shengkang Chen, Shivika Singh, Shivangi Deo 等NeurIPS 2025 · 被引用 1 次
- Self-Organized Polynomial-Time Coordination GraphsQianlan Yang, Weijun Dong, Zhizhou Ren, Jianhao Wang 等ICML 2022 · 被引用 20 次
- Learning NP-Hard Multi-Agent Assignment Planning using GNN: Inference on a Random Graph and Provable Auction-Fitted Q-learningHyunwook Kang, Taehwan Kwon, Jinkyoo Park, James R. MorrisonNeurIPS 2022 · 被引用 4 次
- Autoregressive Policy Optimization for Constrained Allocation TasksDavid Winkel, Niklas Strauß, Maximilian Bernhard, Zongyue Li 等NeurIPS 2024 · 被引用 2 次
- Intersectional Fairness in Reinforcement Learning with Large State and Constraint SpacesEric Eaton, Marcel Hussing, Michael Kearns, Aaron Roth 等ICML 2025
