Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach
Johan Peralez, Aurélien Delage, Jacopo Castellini, Rafael F. Cunha, Jilles Steeve Dibangoye
摘要
The centralized training for decentralized execution paradigm emerged as the state-of-the-art approach to ϵ-optimally solving decentralized partially observable Markov decision processes. However, scalability remains a significant issue. This paper presents a novel and more scalable alternative, namely the sequential-move centralized training for decentralized execution. This paradigm further pushes the applicability of the Bellman’s principle of optimality, raising three new properties. First, it allows a central planner to reason upon sufficient sequential-move statistics instead of prior simultaneous-move ones. Next, it proves that ϵ-optimal value functions are piecewise linear and convex in such sufficient sequential-move statistics. Finally, it drops the complexity of the backup operators from double exponential to polynomial at the expense of longer planning horizons. Besides, it makes it easy to use single-agent methods, e.g., SARSA algorithm enhanced with these findings, while still preserving convergence guarantees. Experiments on two- as well as many-agent domains from the literature against ϵ-optimal simultaneous-move solvers confirm the superiority of our novel approach. This paradigm opens the door for efficient planning and reinforcement learning methods for multi-agent systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 等KDD 2021 · 被引用 49 次
- Solving Common-Payoff Games with Approximate Policy IterationSamuel Sokota, Edward Lockhart, Finbarr Timbers, Elnaz Davoodi 等AAAI 2021 · 被引用 22 次
- Optimally Solving Two-Agent Decentralized POMDPs Under One-Sided Information SharingYuxuan Xie, Jilles Dibangoye, Olivier BuffetICML 2020 · 被引用 15 次
相关 Paper
- Solving Hierarchical Information-Sharing Dec-POMDPs: An Extensive-Form Game ApproachJohan Peralez, Aurélien Delage, Olivier Buffet, Jilles Steeve DibangoyeICML 2024 · 被引用 5 次
- Decentralized TD Tracking with Linear Function Approximation and its Finite-Time AnalysisGang Wang, Songtao Lu, Georgios B. Giannakis, Gerald Tesauro 等NeurIPS 2020 · 被引用 30 次
- Multi-agent active perception with prediction rewardsMikko Lauri, Frans A. OliehoekNeurIPS 2020 · 被引用 13 次
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Asynchronous Actor-Critic for Multi-Agent Reinforcement LearningYuchen Xiao, Weihao Tan, Christopher AmatoNeurIPS 2022 · 被引用 35 次
