Communicating via Markov Decision Processes
Samuel Sokota, Christian A. Schröder de Witt, Maximilian Igl, Luisa M. Zintgraf, Philip H. S. Torr, Martin Strohmeier, J. Zico Kolter, Shimon Whiteson, Jakob N. Foerster
Abstract
We consider the problem of communicating exogenous information by means of Markov decision process trajectories. This setting, which we call a Markov coding game (MCG), generalizes both source coding and a large class of referential games. MCGs also isolate a problem that is important in decentralized control settings in which cheap-talk is not available—namely, they require balancing communication with the associated cost of communicating. We contribute a theoretically grounded approach to MCGs based on maximum entropy reinforcement learning and minimum entropy coupling that we call MEME. Due to recent breakthroughs in approximation algorithms for minimum entropy coupling, MEME is not merely a theoretical algorithm, but can be applied to prac-tical settings. Empirically, we show both that MEME is able to outperform a strong baseline on small MCGs and that MEME is able to achieve strong performance on extremely large MCGs. To the latter point, we demonstrate that MEME is able to losslessly communicate binary images via trajectories of Cartpole and Pong, while simultaneously achieving the maximal or near maximal expected returns, and that it is even capable of performing well in the presence of actuator noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8ca6af0-3ef7-4c85-a860-64c6f10bb3ddCited by top-tier papers5
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- Perfectly Secure Steganography Using Minimum Entropy CouplingChristian Schröder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Nicolaus Foerster et al.ICLR 2023 · 12 citations
- Minimum Entropy Coupling with BottleneckM. Reza Ebrahimi, Jun Chen, Ashish KhistiNeurIPS 2024 · 11 citations
- Learning to Match Unpaired Data with Minimum Entropy CouplingMustapha Bounoua, Giulio Franzese, Pietro MichiardiICML 2025
- Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement LearningYat Long Lo, Christian Schröder de Witt, Samuel Sokota, Jakob Nicolaus Foerster et al.ICLR 2023
Builds on3
- Learning Efficient Multi-agent Communication: An Information Bottleneck ApproachRundong Wang, Xu He, Runsheng Yu, Wei Qiu et al.ICML 2020 · 133 citations
- Learning Agent Communication under Limited Bandwidth by Message PruningHangyu Mao, Zhengchao Zhang, Zhen Xiao, Zhibo Gong et al.AAAI 2020 · 110 citations
- Inference-Based Deterministic Messaging For Multi-Agent CommunicationVarun Bhatt, Michael BuroAAAI 2021 · 5 citations
Related papers
- Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision ProcessesHaotian Wu, Gongpu Chen, Deniz GündüzICLR 2025
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 44 citations
- Maximum Entropy Heterogeneous-Agent Reinforcement LearningJiarong Liu, Yifan Zhong, Siyi Hu, Haobo Fu et al.ICLR 2024 · 27 citations
- Learning Multi-Agent Communication with Contrastive LearningYat Long Lo, Biswa Sengupta, Jakob Nicolaus Foerster, Michael NoukhovitchICLR 2024 · 11 citations
- Uncoupled and Convergent Learning in Two-Player Zero-Sum Markov Games with Bandit FeedbackYang Cai, Haipeng Luo, Chen-Yu Wei, Weiqiang ZhengNeurIPS 2023 · 31 citations
