Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming
Sachin G. Konan, Esmaeil Seraj, Matthew C. Gombolay
Abstract
Information sharing is key in building team cognition and enables coordination and cooperation. High-performing human teams also benefit from acting strategically with hierarchical levels of iterated communication and rationalizability, meaning a human agent can reason about the actions of their teammates in their decision-making. Yet, the majority of prior work in Multi-Agent Reinforcement Learning (MARL) does not support iterated rationalizability and only encourage inter-agent communication, resulting in a suboptimal equilibrium cooperation strategy. In this work, we show that reformulating an agent's policy to be conditional on the policies of its neighboring teammates inherently maximizes Mutual Information (MI) lower-bound when optimizing under Policy Gradient (PG). Building on the idea of decision-making under bounded rationality and cognitive hierarchy theory, we show that our modified PG approach not only maximizes local agent rewards but also implicitly reasons about MI between agents without the need for any explicit ad-hoc regularization terms. Our approach, InfoPG, outperforms baselines in learning emergent collaborative behaviors and sets the state-of-the-art in decentralized cooperative MARL tasks. Our experiments validate the utility of InfoPG by achieving higher sample efficiency and significantly larger cumulative reward in several complex cooperative multi-agent domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8997b48c-833f-4564-8752-87b7b3c6e34cCited by top-tier papers8
- I2Q: A Fully Decentralized Q-Learning AlgorithmJiechuan Jiang, Zongqing LuNeurIPS 2022 · 34 citations
- Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot TeamsEsmaeil Seraj, Jerry Xiong, Mariah Schrum, Matthew C. GombolayNeurIPS 2023 · 12 citations
- Mutual-Information Regularized Multi-Agent Policy IterationJiangxing Wang, Deheng Ye, Zongqing LuNeurIPS 2023 · 9 citations
- DM²: Decentralized Multi-Agent Reinforcement Learning via Distribution MatchingCaroline Wang, Ishan Durugkar, Elad Liebman, Peter StoneAAAI 2023 · 8 citations
- Multi-Agent Coordination via Multi-Level CommunicationGang Ding, Zeyuan Liu, Zhirui Fang, Kefan Su et al.NeurIPS 2024 · 3 citations
Builds on5
- Multi-Agent Game Abstraction via Graph Attention Neural NetworkYong Liu, Weixun Wang, Yujing Hu, Jianye Hao et al.AAAI 2020 · 316 citations
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 156 citations
- The Utility of Explainable AI in Ad Hoc Human-Machine TeamingRohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen et al.NeurIPS 2021 · 103 citations
- CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement LearningJiachen Yang, Alireza Nakhaei, David Isele, Kikuo Fujimura et al.ICLR 2020 · 86 citations
- Byzantine-Resilient Non-Convex Stochastic Gradient DescentZeyuan Allen-Zhu, Faeze Ebrahimianghazani, Jerry Li, Dan AlistarhICLR 2021 · 18 citations
Related papers
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
- Multi-Agent Incentive Communication via Decentralized Teammate ModelingLei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang et al.AAAI 2022 · 104 citations
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar et al.AAAI 2021 · 93 citations
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang et al.NeurIPS 2022 · 65 citations
- Think How Your Teammates Think: Active Inference Can Benefit Decentralized ExecutionHao Wu, Shoucheng Song, Chang Yao, Sheng Han et al.AAAI 2026 · 1 citation
