Episodic Multi-agent Reinforcement Learning with Curiosity-driven Exploration
Lulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He, Yujing Hu, Yingfeng Chen, Changjie Fan, Yang Gao, Chongjie Zhang
摘要
Efficient exploration in deep cooperative multi-agent reinforcement learning (MARL) still remains challenging in complex coordination problems. In this paper, we introduce a novel Episodic Multi-agent reinforcement learning with Curiosity-driven exploration, called EMC. We leverage an insight of popular factorized MARL algorithms that the "induced" individual Q-values, i.e., the individual utility functions used for local execution, are the embeddings of local actionobservation histories, and can capture the interaction between agents due to reward backpropagation during centralized training. Therefore, we use prediction errors of individual Q-values as intrinsic rewards for coordinated exploration and utilize episodic memory to exploit explored informative experience to boost policy training. As the dynamics of an agent's individual Q-value function captures the novelty of states and the influence from other agents, our intrinsic reward can induce coordinated exploration to new or promising states. We illustrate the advantages of our method by didactic examples, and demonstrate its significant outperformance over state-of-the-art MARL baselines on challenging tasks in the StarCraft II micromanagement benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Heterogeneous Skill Learning for Multi-agent TasksYuntao Liu, Yuan Li, Xinhai Xu, Yong Dou 等NeurIPS 2022 · 被引用 33 次
- FoX: Formation-Aware Exploration in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Junghyuk Yeom, Seungyul HanAAAI 2024 · 被引用 22 次
- Efficient Model-based Multi-agent Reinforcement Learning via Optimistic Equilibrium ComputationPier Giuseppe Sessa, Maryam Kamgarpour, Andreas KrauseICML 2022 · 被引用 22 次
- An Adaptive Entropy-Regularization Framework for Multi-Agent Reinforcement LearningWoojun Kim, Youngchul SungICML 2023 · 被引用 20 次
- Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement LearningHyungho Na, Yunkyeong Seo, Il-Chul MoonICLR 2024 · 被引用 13 次
它引用的顶会 Paper6
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- Influence-Based Multi-Agent ExplorationTonghan Wang, Jianhao Wang, Yi Wu, Chongjie ZhangICLR 2020 · 被引用 156 次
相关 Paper
- MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay BufferJeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul SungICML 2022 · 被引用 53 次
- Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Xingchen Li 等NeurIPS 2023 · 被引用 7 次
- Settling Decentralized Multi-Agent Coordinated Exploration by Novelty SharingHaobin Jiang, Ziluo Ding, Zongqing LuAAAI 2024 · 被引用 12 次
- LIGS: Learnable Intrinsic-Reward Generation Selection for Multi-Agent LearningDavid Henry Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等ICLR 2022 · 被引用 20 次
- Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual CalibrationYiyuan Pan, Zhe Liu, Hesheng WangNeurIPS 2025 · 被引用 10 次
