Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming
Sachin G. Konan, Esmaeil Seraj, Matthew C. Gombolay
摘要
Information sharing is key in building team cognition and enables coordination and cooperation. High-performing human teams also benefit from acting strategically with hierarchical levels of iterated communication and rationalizability, meaning a human agent can reason about the actions of their teammates in their decision-making. Yet, the majority of prior work in Multi-Agent Reinforcement Learning (MARL) does not support iterated rationalizability and only encourage inter-agent communication, resulting in a suboptimal equilibrium cooperation strategy. In this work, we show that reformulating an agent's policy to be conditional on the policies of its neighboring teammates inherently maximizes Mutual Information (MI) lower-bound when optimizing under Policy Gradient (PG). Building on the idea of decision-making under bounded rationality and cognitive hierarchy theory, we show that our modified PG approach not only maximizes local agent rewards but also implicitly reasons about MI between agents without the need for any explicit ad-hoc regularization terms. Our approach, InfoPG, outperforms baselines in learning emergent collaborative behaviors and sets the state-of-the-art in decentralized cooperative MARL tasks. Our experiments validate the utility of InfoPG by achieving higher sample efficiency and significantly larger cumulative reward in several complex cooperative multi-agent domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- I2Q: A Fully Decentralized Q-Learning AlgorithmJiechuan Jiang, Zongqing LuNeurIPS 2022 · 被引用 34 次
- Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot TeamsEsmaeil Seraj, Jerry Xiong, Mariah Schrum, Matthew C. GombolayNeurIPS 2023 · 被引用 12 次
- Mutual-Information Regularized Multi-Agent Policy IterationJiangxing Wang, Deheng Ye, Zongqing LuNeurIPS 2023 · 被引用 9 次
- DM²: Decentralized Multi-Agent Reinforcement Learning via Distribution MatchingCaroline Wang, Ishan Durugkar, Elad Liebman, Peter StoneAAAI 2023 · 被引用 8 次
- Multi-Agent Coordination via Multi-Level CommunicationGang Ding, Zeyuan Liu, Zhirui Fang, Kefan Su 等NeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper5
- Multi-Agent Game Abstraction via Graph Attention Neural NetworkYong Liu, Weixun Wang, Yujing Hu, Jianye Hao 等AAAI 2020 · 被引用 316 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
- The Utility of Explainable AI in Ad Hoc Human-Machine TeamingRohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen 等NeurIPS 2021 · 被引用 103 次
- CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement LearningJiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura 等ICLR 2020 · 被引用 86 次
- Byzantine-Resilient Non-Convex Stochastic Gradient DescentZeyuan Allen-Zhu, Faeze Ebrahimianghazani, Jerry Li, Dan AlistarhICLR 2021 · 被引用 18 次
相关 Paper
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang 等ICML 2022 · 被引用 49 次
- Multi-Agent Incentive Communication via Decentralized Teammate ModelingLei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang 等AAAI 2022 · 被引用 104 次
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar 等AAAI 2021 · 被引用 93 次
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang 等NeurIPS 2022 · 被引用 65 次
- Think How Your Teammates Think: Active Inference Can Benefit Decentralized ExecutionHao Wu, Shoucheng Song, Chang Yao, Sheng Han 等AAAI 2026 · 被引用 1 次
