Learning Nearly Decomposable Value Functions Via Communication Minimization
Tonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie Zhang
摘要
Reinforcement learning encounters major challenges in multi-agent settings, such as scalability and non-stationarity. Recently, value function factorization learning emerges as a promising way to address these challenges in collaborative multiagent systems. However, existing methods have been focusing on learning fully decentralized value functions, which are not efficient for tasks requiring communication. To address this limitation, this paper presents a novel framework for learning nearly decomposable Q-functions (NDQ) via communication minimization, with which agents act on their own most of the time but occasionally send messages to other agents in order for effective coordination. This framework hybridizes value function factorization learning and communication learning by introducing two information-theoretic regularizers. These regularizers are maximizing mutual information between agents' action selection and communication messages while minimizing the entropy of messages between agents. We show how to optimize these regularizers in a way that is easily integrated with existing value function factorization methods such as QMIX. Finally, we demonstrate that, on the StarCraft unit micromanagement benchmark, our framework significantly outperforms baseline methods and allows us to cut off more than 80% of communication without sacrificing the performance. The videos of our experiments are available at https://sites.google.com/view/ndq .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Learning Individually Inferred Communication for Multi-Agent CooperationZiluo Ding, Tiejun Huang, Zongqing LuNeurIPS 2020 · 被引用 146 次
- Multi-Agent Incentive Communication via Decentralized Teammate ModelingLei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang 等AAAI 2022 · 被引用 104 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang 等NeurIPS 2022 · 被引用 65 次
相关 Paper
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 被引用 140 次
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-LearningWei-Fang Sun, Cheng-Kuang Lee, Chun-Yi LeeICML 2021 · 被引用 56 次
- Learning Efficient and Robust Multi-Agent Communication via Graph Information BottleneckShifei Ding, Wei Du, Ling Ding, Lili Guo 等AAAI 2024 · 被引用 12 次
- ConcaveQ: Non-monotonic Value Function Factorization via Concave Representations in Deep Multi-Agent Reinforcement LearningHuiqun Li, Hanhan Zhou, Yifei Zou, Dongxiao Yu 等AAAI 2024 · 被引用 18 次
- PAC: Assisted Value Factorization with Counterfactual Predictions in Multi-Agent Reinforcement LearningHanhan Zhou, Tian Lan, Vaneet AggarwalNeurIPS 2022 · 被引用 47 次
