ELIGN: Expectation Alignment as a Multi-Agent Intrinsic Reward
Zixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein, Ranjay Krishna
摘要
Modern multi-agent reinforcement learning frameworks rely on centralized training and reward shaping to perform well. However, centralized training and dense rewards are not readily available in the real world. Current multi-agent algorithms struggle to learn in the alternative setup of decentralized training or sparse rewards. To address these issues, we propose a self-supervised intrinsic reward ELIGN - expectation alignment - inspired by the self-organization principle in Zoology. Similar to how animals collaborate in a decentralized manner with those in their vicinity, agents trained with expectation alignment learn behaviors that match their neighbors' expectations. This allows the agents to learn collaborative behaviors without any external reward or centralized training. We demonstrate the efficacy of our approach across 6 tasks in the multi-agent particle and the complex Google Research football environments, comparing ELIGN to sparse and curiosity-based intrinsic rewards. When the number of agents increases, ELIGN scales well in all multi-agent tasks except for one where agents have different capabilities. We show that agent coordination improves through expectation alignment because agents learn to divide tasks amongst themselves, break coordination symmetries, and confuse adversaries. These results identify tasks where expectation alignment is a more useful strategy than curiosity-driven exploration for multi-agent coordination, enabling agents to do zero-shot coordination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement LearningJunkai Zhang, Yifan Zhang, Xi Sheryl Zhang, Yifan Zang 等AAAI 2024 · 被引用 9 次
- Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement LearningHao Ma, Shijie Wang, Zhiqiang Pu, Siyao Zhao 等AAAI 2025 · 被引用 1 次
- GeoExplorer: Active Geo-Localization with Curiosity-Driven ExplorationLi Mi, Manon Béchaz, Zeming Chen, Antoine Bosselut 等ICCV 2025
它引用的顶会 Paper8
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 被引用 61 次
相关 Paper
- Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement LearningRui Tang, Biao Luo, Yongzheng CuiAAAI 2026
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai 等NeurIPS 2020 · 被引用 25 次
- Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement LearningXinran Li, Zifan Liu, Shibo Chen, Jun ZhangICML 2024 · 被引用 11 次
- Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksPei Xu, Junge Zhang, Qiyue Yin, Chao Yu 等AAAI 2023 · 被引用 15 次
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang 等ICML 2022 · 被引用 49 次
