Surprise Minimizing Multi-Agent Learning with Energy-based Models
Karush Suri, Xiao Qi Shi, Konstantinos N. Plataniotis, Yuri A. Lawryshyn
摘要
Multi-Agent Reinforcement Learning (MARL) has demonstrated significant success by virtue of collaboration across agents. Recent work, on the other hand, introduces surprise which quantifies the degree of change in an agent's environment. Surprise-based learning has received significant attention in the case of single-agent entropic settings but remains an open problem for fast-paced dynamics in multi-agent scenarios. A potential alternative to address surprise may be realized through the lens of free-energy minimization. We explore surprise minimization in multi-agent learning by utilizing the free energy across all agents in a multi-agent system. A temporal Energy-Based Model (EBM) represents an estimate of surprise which is minimized over the joint agent distribution. Our formulation of the EBM is theoretically akin to the minimum conjugate entropy objective and highlights suitable convergence towards minimum surprising states. We further validate our theoretical claims in an empirical study of multi-agent tasks demanding collaboration in the presence of fast-paced dynamics. Our implementation and agent videos are available at the Project Webpage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao 等NeurIPS 2021 · 被引用 224 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- RODE: Learning Roles to Decompose Multi-Agent TasksTonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng 等ICLR 2021 · 被引用 60 次
相关 Paper
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable EnvironmentsGlen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart 等ICLR 2021 · 被引用 12 次
- A Mixture Of Surprises for Unsupervised Reinforcement LearningAndrew Zhao, Matthieu Gaetan Lin, Yangguang Li, Yong-Jin Liu 等NeurIPS 2022 · 被引用 15 次
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 被引用 58 次
- ELMA: Energy-Based Learning for Multi-Agent Activity ForecastingYu-Ke Li, Pin Wang, Lixiong Chen, Zheng Wang 等AAAI 2022 · 被引用 8 次
- TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task OptimizationWonhyeok Choi, Kyumin Hwang, Jihun Park, Kyoungmin Lee 等CVPR 2026
