RGMComm: Return Gap Minimization via Discrete Communications in Multi-Agent Reinforcement Learning
Jingdi Chen, Tian Lan, Carlee Joe-Wong
Abstract
Communication is crucial for solving cooperative Multi-Agent Reinforcement Learning tasks in partially observable Markov Decision Processes. Existing works often rely on black-box methods to encode local information/features into messages shared with other agents, leading to the generation of continuous messages with high communication overhead and poor interpretability. Prior attempts at discrete communication methods generate one-hot vectors trained as part of agents' actions and use the Gumbel softmax operation for calculating message gradients, which are all heuristic designs that do not provide any quantitative guarantees on the expected return. This paper establishes an upper bound on the return gap between an ideal policy with full observability and an optimal partially observable policy with discrete communication. This result enables us to recast multi-agent communication into a novel online clustering problem over the local observations at each agent, with messages as cluster labels and the upper bound on the return gap as clustering loss. To minimize the return gap, we propose the Return-Gap-Minimization Communication (RGMComm) algorithm, which is a surprisingly simple design of discrete message generation functions and is integrated with reinforcement learning through the utilization of a novel Regularized Information Maximization loss function, which incorporates cosine-distance as the clustering metric. Evaluations show that RGMComm significantly outperforms state-of-the-art multi-agent communication baselines and can achieve nearly optimal returns with few-bit messages that are naturally interpretable. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6edaeace-9d57-499c-ab03-7730665ca9d3Cited by top-tier papers4
- Conspirator: SmartNIC-Aided Control Plane for Distributed ML WorkloadsYunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao et al.USENIX ATC 2024 · 13 citations
- RGMDT: Return-Gap-Minimizing Decision Tree Extraction in Non-Euclidean Metric SpaceJingdi Chen, Hanhan Zhou, Yongsheng Mei, Carlee Joe-Wong et al.NeurIPS 2024 · 2 citations
- Learning to Communicate Through Implicit Communication ChannelsHan Wang, Binbin Chen, Tieying Zhang, Baoxiang WangICLR 2025
- Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision ProcessesHaotian Wu, Gongpu Chen, Deniz GündüzICLR 2025
Builds on4
- Learning Nearly Decomposable Value Functions Via Communication MinimizationTonghan Wang, Jianhao Wang, Chongyi Zheng, Chongjie ZhangICLR 2020 · 170 citations
- FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement LearningTianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie et al.ICML 2021 · 88 citations
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 42 citations
- Emergent Discrete Communication in Semantic SpacesMycal Tucker, Huao Li, Siddharth Agrawal, Dana Hughes et al.NeurIPS 2021 · 34 citations
Related papers
- Emergent Quantized CommunicationBoaz Carmeli, Ron Meir, Yonatan BelinkovAAAI 2023 · 10 citations
- Learning Efficient and Interpretable Multi-Agent CommunicationWei Du, Benyu Wu, Yuqing Sun, Wei Guo et al.ICLR 2026
- Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement LearningYat Long Lo, Christian Schröder de Witt, Samuel Sokota, Jakob Nicolaus Foerster et al.ICLR 2023
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang et al.NeurIPS 2022 · 65 citations
- Multi-Agent Incentive Communication via Decentralized Teammate ModelingLei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang et al.AAAI 2022 · 104 citations
