ELIGN: Expectation Alignment as a Multi-Agent Intrinsic Reward
Zixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein, Ranjay Krishna
Abstract
Modern multi-agent reinforcement learning frameworks rely on centralized training and reward shaping to perform well. However, centralized training and dense rewards are not readily available in the real world. Current multi-agent algorithms struggle to learn in the alternative setup of decentralized training or sparse rewards. To address these issues, we propose a self-supervised intrinsic reward ELIGN - expectation alignment - inspired by the self-organization principle in Zoology. Similar to how animals collaborate in a decentralized manner with those in their vicinity, agents trained with expectation alignment learn behaviors that match their neighbors' expectations. This allows the agents to learn collaborative behaviors without any external reward or centralized training. We demonstrate the efficacy of our approach across 6 tasks in the multi-agent particle and the complex Google Research football environments, comparing ELIGN to sparse and curiosity-based intrinsic rewards. When the number of agents increases, ELIGN scales well in all multi-agent tasks except for one where agents have different capabilities. We show that agent coordination improves through expectation alignment because agents learn to divide tasks amongst themselves, break coordination symmetries, and confuse adversaries. These results identify tasks where expectation alignment is a more useful strategy than curiosity-driven exploration for multi-agent coordination, enabling agents to do zero-shot coordination.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5646c21c-52d3-4822-8628-d73ace45d4d4Cited by top-tier papers3
- Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement LearningJunkai Zhang, Yifan Zhang, Xi Sheryl Zhang, Yifan Zang et al.AAAI 2024 · 9 citations
- Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement LearningHao Ma, Shijie Wang, Zhiqiang Pu, Siyao Zhao et al.AAAI 2025 · 1 citation
- GeoExplorer: Active Geo-Localization with Curiosity-Driven ExplorationLi Mi, Manon Béchaz, Zeming Chen, Antoine Bosselut et al.ICCV 2025
Builds on8
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 156 citations
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 61 citations
Related papers
- Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement LearningRui Tang, Biao Luo, Yongzheng CuiAAAI 2026
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai et al.NeurIPS 2020 · 25 citations
- Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement LearningXinran Li, Zifan Liu, Shibo Chen, Jun ZhangICML 2024 · 11 citations
- Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksPei Xu, Junge Zhang, Qiyue Yin, Chao Yu et al.AAAI 2023 · 15 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
