Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment
Chen Zhang, Qiang He, Yuan Zhou, Elvis S. Liu, Hong Wang, Jian Zhao, Yang Wang
Abstract
Deep Reinforcement Learning (DRL) agents have demonstrated impressive success in a wide range of game genres. However, existing research primarily focuses on optimizing DRL competence rather than addressing the challenge of prolonged player interaction. In this paper, we propose a practical DRL agent system for fighting games named Sh=ukai, which has been successfully deployed to Naruto Mobile, a popular fighting game with over 100 million registered users. Sh=ukai quantifies the state to enhance generalizability, introducing Heterogeneous League Training (HELT) to achieve balanced competence, generalizability, and training efficiency. Furthermore, Sh=ukai implements specific rewards to align the agent's behavior with human expectations. Sh=ukai's ability to generalize is demonstrated by its consistent competence across all characters, even though it was trained on only 13% of them. Additionally, HELT exhibits a remarkable 22% improvement in sample efficiency. Sh=ukai serves as a valuable training partner for players in Naruto Mobile, enabling them to enhance their abilities and skills.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc3d8e47-19a8-4c28-8476-624b80d296faCited by top-tier papers1
Ask how each one uses itBuilds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi et al.AAAI 2020 · 395 citations
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen et al.NeurIPS 2020 · 225 citations
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
- Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningSiming Lan, Rui Zhang, Qi Yi, Jiaming Guo et al.NeurIPS 2023 · 18 citations
Related papers
- Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless ComputingHanfei Yu, Jian Li, Yang Hua, Xu Yuan et al.AAAI 2024 · 8 citations
- Nitro: Boosting Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Jacob Carter, Hao Wang, Devesh Tiwari et al.VLDB 2025 · 3 citations
- Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PCTyler Clark, Mark Towers, Christine Evers, Jonathon HareICML 2025
- DRMD: Deep Reinforcement Learning for Malware Detection Under Concept DriftShae McFadden, Myles Foley, Mario D'Onghia, Chris Hicks et al.AAAI 2026 · 7 citations
- Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainYiming Gao, Feiyu Liu, Liang Wang, Dehua Zheng et al.ICLR 2024 · 5 citations
