Correcting experience replay for multi-agent communication
Sanjeevan Ahilan, Peter Dayan
摘要
We consider the problem of learning to communicate using multi-agent reinforcement learning (MARL). A common approach is to learn off-policy, using data sampled from a replay buffer. However, messages received in the past may not accurately reflect the current communication policy of each agent, and this complicates learning. We therefore introduce a 'communication correction' which accounts for the non-stationarity of observed communication induced by multi-agent learning. It works by relabelling the received message to make it likely under the communicator's current policy, and thus be a better reflection of the receiver's current environment. To account for cases in which agents are both senders and receivers, we introduce an ordered relabelling scheme. Our correction is computationally efficient and can be integrated with a range of off-policy algorithms. We find in our experiments that it substantially improves the ability of communicating MARL systems to learn across a variety of cooperative and competitive tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Efficient Multi-agent Communication via Self-supervised Information AggregationCong Guan, Feng Chen, Lei Yuan, Chenghe Wang 等NeurIPS 2022 · 被引用 65 次
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 被引用 2 次
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 被引用 42 次
- PMAC: Personalized Multi-Agent CommunicationXiangrui Meng, Ying TanAAAI 2024 · 被引用 7 次
- Robust Communicative Multi-Agent Reinforcement Learning with Active DefenseLebin Yu, Yunbo Qiu, Quanming Yao, Yuan Shen 等AAAI 2024 · 被引用 11 次
