Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach
Johan Peralez, Aurélien Delage, Jacopo Castellini, Rafael F. Cunha, Jilles Steeve Dibangoye
Abstract
The centralized training for decentralized execution paradigm emerged as the state-of-the-art approach to ϵ-optimally solving decentralized partially observable Markov decision processes. However, scalability remains a significant issue. This paper presents a novel and more scalable alternative, namely the sequential-move centralized training for decentralized execution. This paradigm further pushes the applicability of the Bellman’s principle of optimality, raising three new properties. First, it allows a central planner to reason upon sufficient sequential-move statistics instead of prior simultaneous-move ones. Next, it proves that ϵ-optimal value functions are piecewise linear and convex in such sufficient sequential-move statistics. Finally, it drops the complexity of the backup operators from double exponential to polynomial at the expense of longer planning horizons. Besides, it makes it easy to use single-agent methods, e.g., SARSA algorithm enhanced with these findings, while still preserving convergence guarantees. Experiments on two- as well as many-agent domains from the literature against ϵ-optimal simultaneous-move solvers confirm the superiority of our novel approach. This paradigm opens the door for efficient planning and reinforcement learning methods for multi-agent systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d973bdd-7646-4b72-ab6e-a920bbcd61f1Cited by top-tier papers1
Ask how each one uses itBuilds on6
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.KDD 2021 · 49 citations
- Solving Common-Payoff Games with Approximate Policy IterationSamuel Sokota, Edward Lockhart, Finbarr Timbers, Elnaz Davoodi et al.AAAI 2021 · 22 citations
- Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information SharingYuxuan Xie, Jilles Dibangoye, Olivier BuffetICML 2020 · 15 citations
Related papers
- Solving Hierarchical Information-Sharing Dec-POMDPs: An Extensive-Form Game ApproachJohan Peralez, Aurélien Delage, Olivier Buffet, Jilles Steeve DibangoyeICML 2024 · 5 citations
- Decentralized TD Tracking with Linear Function Approximation and its Finite-Time AnalysisGang Wang, Songtao Lu, Georgios B. Giannakis, Gerald Tesauro et al.NeurIPS 2020 · 30 citations
- Multi-agent active perception with prediction rewardsMikko Lauri, Frans A. OliehoekNeurIPS 2020 · 13 citations
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Asynchronous Actor-Critic for Multi-Agent Reinforcement LearningYuchen Xiao, Weihao Tan, Christopher AmatoNeurIPS 2022 · 35 citations
