Multi-agent Markov Entanglement
Shuze Chen, Tianyi Peng
Abstract
Value decomposition has long been a fundamental technique in multi-agent dynamic programming and reinforcement learning (RL). Specifically, the value function of a global state is often approximated as the sum of local functions: . This approach traces back to the index policy in restless multi-armed bandit problems and has found various applications in modern RL systems. However, the theoretical justification for why this decomposition works so effectively remains underexplored. In this paper, we uncover the underlying mathematical structure that enables value decomposition. We demonstrate that a multi-agent Markov decision process (MDP) permits value decomposition if and only if its transition matrix is not"entangled"-- a concept analogous to quantum entanglement in quantum physics. Drawing inspiration from how physicists measure quantum entanglement, we introduce how to measure the"Markov entanglement"for multi-agent MDPs and show that this measure can be used to bound the decomposition error in general multi-agent MDPs. Using the concept of Markov entanglement, we proved that a widely-used class of index policies is weakly entangled and enjoys a sublinear scale of decomposition error for -agent systems. Finally, we show how Markov entanglement can be efficiently estimated in practice, providing practitioners with an empirical proxy for the quality of value decomposition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66671596-3a39-424b-8d2b-a093781b9dd0Builds on8
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 156 citations
- NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RLKhaled Nakhleh, Santosh Ganji, Ping-Chun Hsieh, I-Hong Hou et al.NeurIPS 2021 · 52 citations
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye et al.NeurIPS 2021 · 50 citations
- Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement LearningYitian Hong, Yaochu Jin, Yang TangNeurIPS 2022 · 40 citations
Related papers
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningJianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun KimNeurIPS 2022 · 50 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu et al.NeurIPS 2022 · 35 citations
- Detecting Influence Structures in Multi-Agent Reinforcement LearningFabian Raoul Pieroth, Katherine E. Fitch, Lenz BelznerICML 2024 · 2 citations
- Value Function Decomposition for Iterative Design of Reinforcement Learning AgentsJames MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno et al.NeurIPS 2022 · 12 citations
