Mastering Complex Control in MOBA Games with Deep Reinforcement Learning
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin
摘要
We study the reinforcement learning problem of complex action control in the Multi-player Online Battle Arena (MOBA) 1v1 games. This problem involves far more complicated state and action spaces than those of traditional 1v1 games, such as Go and Atari series, which makes it very difficult to search any policies with human-level performance. In this paper, we present a deep reinforcement learning framework to tackle this problem from the perspectives of both system and algorithm. Our system is of low coupling and high scalability, which enables efficient explorations at large scale. Our algorithm includes several novel strategies, including control dependency decoupling, action mask, target attention, and dualclip PPO, with which our proposed actor-critic network can be effectively trained in our system. Tested on the MOBA game Honor of Kings, our AI agent, called Tencent Solo, can defeat top professional human players in full 1v1 games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper50
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen 等NeurIPS 2020 · 被引用 225 次
- DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement LearningDaochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang 等ICML 2021 · 被引用 150 次
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng 等NeurIPS 2021 · 被引用 133 次
- Sample More to Think Less: Group Filtered Policy Optimization for Concise ReasoningVaishnavi Shrivastava, Ahmed Hassan Awadallah, Vidhisha Balachandran, Shivam Garg 等ICLR 2026 · 被引用 85 次
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
相关 Paper
- Learning Diverse Policies in MOBA Games via Macro-GoalsYiming Gao, Bei Shi, Xueying Du, Liang Wang 等NeurIPS 2021 · 被引用 17 次
- Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication PerspectiveYiming Gao, Feiyu Liu, Liang Wang, Zhenjie Lian 等ICLR 2023 · 被引用 2 次
- Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainYiming Gao, Feiyu Liu, Liang Wang, Dehua Zheng 等ICLR 2024 · 被引用 5 次
- SCC: an efficient deep reinforcement learning agent mastering the game of StarCraft IIXiangjun Wang, Junxiao Song, Penghui Qi, Peng Peng 等ICML 2021 · 被引用 50 次
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
