Surprise Minimizing Multi-Agent Learning with Energy-based Models
Karush Suri, Xiao Qi Shi, Konstantinos N. Plataniotis, Yuri A. Lawryshyn
Abstract
Multi-Agent Reinforcement Learning (MARL) has demonstrated significant success by virtue of collaboration across agents. Recent work, on the other hand, introduces surprise which quantifies the degree of change in an agent's environment. Surprise-based learning has received significant attention in the case of single-agent entropic settings but remains an open problem for fast-paced dynamics in multi-agent scenarios. A potential alternative to address surprise may be realized through the lens of free-energy minimization. We explore surprise minimization in multi-agent learning by utilizing the free energy across all agents in a multi-agent system. A temporal Energy-Based Model (EBM) represents an estimate of surprise which is minimized over the joint agent distribution. Our formulation of the EBM is theoretically akin to the minimum conjugate entropy objective and highlights suitable convergence towards minimum surprising states. We further validate our theoretical claims in an empirical study of multi-agent tasks demanding collaboration in the presence of fast-paced dynamics. Our implementation and agent videos are available at the Project Webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a96f548c-acac-4d8e-8435-25dca642f2a7Builds on5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao et al.NeurIPS 2021 · 224 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- RODE: Learning Roles to Decompose Multi-Agent TasksTonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng et al.ICLR 2021 · 60 citations
Related papers
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable EnvironmentsGlen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart et al.ICLR 2021 · 12 citations
- A Mixture Of Surprises for Unsupervised Reinforcement LearningAndrew Zhao, Matthieu Gaetan Lin, Yangguang Li, Yong-Jin Liu et al.NeurIPS 2022 · 15 citations
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 58 citations
- ELMA: Energy-Based Learning for Multi-Agent Activity ForecastingYu-Ke Li, Pin Wang, Lixiong Chen, Zheng Wang et al.AAAI 2022 · 8 citations
- TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task OptimizationWonhyeok Choi, Kyumin Hwang, Jihun Park, Kyoungmin Lee et al.CVPR 2026
