Towards Playing Full MOBA Games with Deep Reinforcement Learning
Deheng Ye, Guibin Chen, Wen Zhang, Sheng Chen, Bo Yuan, Bo Liu, Jia Chen, Zhao Liu, Fuhao Qiu, Hongsheng Yu, Yinyuting Yin, Bei Shi
摘要
MOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handling the raw game complexity caused by the explosion of agent combinations, i.e., lineups, when expanding the hero pool in case that OpenAI's Dota AI limits the play to a pool of only 17 heroes. As a result, full MOBA games without restrictions are far from being mastered by any existing AI system. In this paper, we propose a MOBA AI learning paradigm that methodologically enables playing full MOBA games with deep reinforcement learning. Specifically, we develop a combination of novel and existing learning techniques, including curriculum self-play learning, policy distillation, off-policy adaption, multi-head value estimation, and Monte-Carlo tree-search, in training and playing a large pool of heroes, meanwhile addressing the scalability issue skillfully. Tested on Honor of Kings, a popular MOBA game, we show how to build superhuman AI agents that can defeat top esports players. The superiority of our AI is demonstrated by the first large-scale performance test of MOBA AI agent in the literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
- DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement LearningDaochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang 等ICML 2021 · 被引用 150 次
- Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting PotJoel Z. Leibo, Edgar A. Duéñez-Guzmán, Alexander Vezhnevets, John P. Agapiou 等ICML 2021 · 被引用 134 次
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade 等NeurIPS 2022 · 被引用 115 次
- Modelling Behavioural Diversity for Learning in Open-Ended GamesNicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Henry Mguni 等ICML 2021 · 被引用 80 次
它引用的顶会 Paper2
相关 Paper
- Learning Diverse Policies in MOBA Games via Macro-GoalsYiming Gao, Bei Shi, Xueying Du, Liang Wang 等NeurIPS 2021 · 被引用 17 次
- Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication PerspectiveYiming Gao, Feiyu Liu, Liang Wang, Zhenjie Lian 等ICLR 2023 · 被引用 2 次
- Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainYiming Gao, Feiyu Liu, Liang Wang, Dehua Zheng 等ICLR 2024 · 被引用 5 次
- Learning to Identify Top Elo Ratings: A Dueling Bandits ApproachXue Yan, Yali Du, Binxin Ru, Jun Wang 等AAAI 2022 · 被引用 9 次
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray 等ICLR 2023 · 被引用 10 次
