Communicating via Markov Decision Processes
Samuel Sokota, Christian A. Schröder de Witt, Maximilian Igl, Luisa M. Zintgraf, Philip H. S. Torr, Martin Strohmeier, J. Zico Kolter, Shimon Whiteson, Jakob N. Foerster
摘要
We consider the problem of communicating exogenous information by means of Markov decision process trajectories. This setting, which we call a Markov coding game (MCG), generalizes both source coding and a large class of referential games. MCGs also isolate a problem that is important in decentralized control settings in which cheap-talk is not available—namely, they require balancing communication with the associated cost of communicating. We contribute a theoretically grounded approach to MCGs based on maximum entropy reinforcement learning and minimum entropy coupling that we call MEME. Due to recent breakthroughs in approximation algorithms for minimum entropy coupling, MEME is not merely a theoretical algorithm, but can be applied to prac-tical settings. Empirically, we show both that MEME is able to outperform a strong baseline on small MCGs and that MEME is able to achieve strong performance on extremely large MCGs. To the latter point, we demonstrate that MEME is able to losslessly communicate binary images via trajectories of Cartpole and Pong, while simultaneously achieving the maximal or near maximal expected returns, and that it is even capable of performing well in the presence of actuator noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- Perfectly Secure Steganography Using Minimum Entropy CouplingChristian Schröder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Nicolaus Foerster 等ICLR 2023 · 被引用 12 次
- Minimum Entropy Coupling with BottleneckM. Reza Ebrahimi, Jun Chen, Ashish KhistiNeurIPS 2024 · 被引用 11 次
- Learning to Match Unpaired Data with Minimum Entropy CouplingMustapha Bounoua, Giulio Franzese, Pietro MichiardiICML 2025
- Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement LearningYat Long Lo, Christian Schröder de Witt, Samuel Sokota, Jakob Nicolaus Foerster 等ICLR 2023
它引用的顶会 Paper3
- Learning Efficient Multi-agent Communication: An Information Bottleneck ApproachRundong Wang, Xu He, Runsheng Yu, Wei Qiu 等ICML 2020 · 被引用 133 次
- Learning Agent Communication under Limited Bandwidth by Message PruningHangyu Mao, Zhengchao Zhang, Zhen Xiao, Zhibo Gong 等AAAI 2020 · 被引用 110 次
- Inference-Based Deterministic Messaging For Multi-Agent CommunicationVarun Bhatt, Michael BuroAAAI 2021 · 被引用 5 次
相关 Paper
- Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision ProcessesHaotian Wu, Gongpu Chen, Deniz GündüzICLR 2025
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 被引用 44 次
- Maximum Entropy Heterogeneous-Agent Reinforcement LearningJiarong Liu, Yifan Zhong, Siyi Hu, Haobo Fu 等ICLR 2024 · 被引用 27 次
- Learning Multi-Agent Communication with Contrastive LearningYat Long Lo, Biswa Sengupta, Jakob Nicolaus Foerster, Michael NoukhovitchICLR 2024 · 被引用 11 次
- Uncoupled and Convergent Learning in Two-Player Zero-Sum Markov Games with Bandit FeedbackYang Cai, Haipeng Luo, Chen-Yu Wei, Weiqiang ZhengNeurIPS 2023 · 被引用 31 次
