Learning to Communicate Implicitly by Actions
Zheng Tian, Shihao Zou, Ian Davies, Tim Warr, Lisheng Wu, Haitham Bou-Ammar, Jun Wang
摘要
In situations where explicit communication is limited, human collaborators act by learning to: (i) infer meaning behind their partner's actions, and (ii) convey private information about the state to their partner implicitly through actions. The first component of this learning process has been well-studied in multi-agent systems, whereas the second — which is equally crucial for successful collaboration — has not. To mimic both components mentioned above, thereby completing the learning process, we introduce a novel algorithm: Policy Belief Learning (PBL). PBL uses a belief module to model the other agent's private information and a policy module to form a distribution over actions informed by the belief module. Furthermore, to encourage communication by actions, we propose a novel auxiliary reward which incentivizes one agent to help its partner to make correct inferences about its private information. The auxiliary reward for communication is integrated into the learning of the policy module. We evaluate our approach on a set of environments including a matrix game, particle environment and the non-competitive bidding problem from contract bridge. We show empirically that this auxiliary reward is effective and easy to generalize. These results demonstrate that our PBL algorithm can produce strong pairs of agents in collaborative games where explicit communication is disabled.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Generalized Beliefs for Cooperative AIDarius Muglich, Luisa M. Zintgraf, Christian A. Schröder de Witt, Shimon Whiteson 等ICML 2022 · 被引用 11 次
- Differentiable Multi-Agent Actor-Critic for Multi-Step Radiology Report SummarizationSanjeev Kumar Karn, Ning Liu, Hinrich Schütze, Oladimeji FarriACL 2022
- CtD: Composition through Decomposition in Emergent CommunicationBoaz Carmeli, Ron Meir, Yonatan BelinkovICLR 2025
- Learning to Communicate Through Implicit Communication ChannelsHan Wang, Binbin Chen, Tieying Zhang, Baoxiang WangICLR 2025
- Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision ProcessesHaotian Wu, Gongpu Chen, Deniz GündüzICLR 2025
相关 Paper
- Learning Individually Inferred Communication for Multi-Agent CooperationZiluo Ding, Tiejun Huang, Zongqing LuNeurIPS 2020 · 被引用 146 次
- Joint Policy Search for Multi-agent Collaboration with Imperfect InformationYuandong Tian, Qucheng Gong, Yu JiangNeurIPS 2020 · 被引用 24 次
- Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial ExplorationAndreas Kontogiannis, Konstantinos Papathanasiou, Yi Shen, Giorgos Stamou 等ICML 2025
- Inference-Based Deterministic Messaging For Multi-Agent CommunicationVarun Bhatt, Michael BuroAAAI 2021 · 被引用 5 次
- Correcting experience replay for multi-agent communicationSanjeevan Ahilan, Peter DayanICLR 2021 · 被引用 3 次
