Multi-agent Markov Entanglement
Shuze Chen, Tianyi Peng
摘要
Value decomposition has long been a fundamental technique in multi-agent dynamic programming and reinforcement learning (RL). Specifically, the value function of a global state is often approximated as the sum of local functions: . This approach traces back to the index policy in restless multi-armed bandit problems and has found various applications in modern RL systems. However, the theoretical justification for why this decomposition works so effectively remains underexplored. In this paper, we uncover the underlying mathematical structure that enables value decomposition. We demonstrate that a multi-agent Markov decision process (MDP) permits value decomposition if and only if its transition matrix is not"entangled"-- a concept analogous to quantum entanglement in quantum physics. Drawing inspiration from how physicists measure quantum entanglement, we introduce how to measure the"Markov entanglement"for multi-agent MDPs and show that this measure can be used to bound the decomposition error in general multi-agent MDPs. Using the concept of Markov entanglement, we proved that a widely-used class of index policies is weakly entangled and enjoys a sublinear scale of decomposition error for -agent systems. Finally, we show how Markov entanglement can be efficiently estimated in practice, providing practitioners with an empirical proxy for the quality of value decomposition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
- NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RLKhaled Nakhleh, Santosh Ganji, Ping-Chun Hsieh, I-Hong Hou 等NeurIPS 2021 · 被引用 52 次
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye 等NeurIPS 2021 · 被引用 50 次
- Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement LearningYitian Hong, Yaochu Jin, Yang TangNeurIPS 2022 · 被引用 40 次
相关 Paper
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningJianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun KimNeurIPS 2022 · 被引用 50 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu 等NeurIPS 2022 · 被引用 35 次
- Detecting Influence Structures in Multi-Agent Reinforcement LearningFabian Raoul Pieroth, Katherine E. Fitch, Lenz BelznerICML 2024 · 被引用 2 次
- Value Function Decomposition for Iterative Design of Reinforcement Learning AgentsJames MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno 等NeurIPS 2022 · 被引用 12 次
