LIGS: Learnable Intrinsic-Reward Generation Selection for Multi-Agent Learning
David Henry Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves, Oliver Slumbers, Feifei Tong, Yang Li, Jiangcheng Zhu, Yaodong Yang, Jun Wang
摘要
Efficient exploration is important for reinforcement learners to achieve high rewards. In multi-agent systems, coordinated exploration and behaviour is critical for agents to jointly achieve optimal outcomes. In this paper, we introduce a new general framework for improving coordination and performance of multi-agent reinforcement learners (MARL). Our framework, named Learnable Intrinsic-Reward Generation Selection algorithm (LIGS) introduces an adaptive learner, Generator that observes the agents and learns to construct intrinsic rewards online that coordinate the agents’ joint exploration and joint behaviour. Using a novel combination of MARL and switching controls, LIGS determines the best states to learn to add intrinsic rewards which leads to a highly efficient learning process. LIGS can subdivide complex tasks making them easier to solve and enables systems of MARL agents to quickly solve environments with sparse rewards. LIGS can seamlessly adopt existing MARL algorithms and, our theory shows that it ensures convergence to policies that deliver higher system performance. We demonstrate its superior performance in challenging tasks in Foraging and StarCraft II.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 等NeurIPS 2022 · 被引用 408 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement LearningJunkai Zhang, Yifan Zhang, Xi Sheryl Zhang, Yifan Zang 等AAAI 2024 · 被引用 9 次
- Open Ad Hoc Teamwork with Cooperative Game TheoryJianhong Wang, Yang Li, Yuan Zhang, Wei Pan 等ICML 2024 · 被引用 5 次
- Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement LearningHao Ma, Shijie Wang, Zhiqiang Pu, Siyao Zhao 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper6
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song 等NeurIPS 2021 · 被引用 216 次
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 被引用 159 次
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu 等NeurIPS 2021 · 被引用 121 次
- Multi-Agent Determinantal Q-LearningYaodong Yang, Ying Wen, Jun Wang, Liheng Chen 等ICML 2020 · 被引用 83 次
- Learning in Nonzero-Sum Stochastic Games with PotentialsDavid Henry Mguni, Yutong Wu, Yali Du, Yaodong Yang 等ICML 2021 · 被引用 51 次
相关 Paper
- MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay BufferJeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul SungICML 2022 · 被引用 53 次
- MANSA: Learning Fast and Slow in Multi-Agent SystemsDavid Henry Mguni, Haojun Chen, Taher Jafferjee, Jianhong Wang 等ICML 2023 · 被引用 4 次
- Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement LearningXinran Li, Zifan Liu, Shibo Chen, Jun ZhangICML 2024 · 被引用 11 次
- Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement LearningBoyin Liu, Zhiqiang Pu, Yi Pan, Jianqiang Yi 等ICML 2023 · 被引用 34 次
- Generative Exploration and ExploitationJiechuan Jiang, Zongqing LuAAAI 2020 · 被引用 6 次
