PerfectDou: Dominating DouDizhu with Perfect Information Distillation
Guan Yang, Minghuan Liu, Weijun Hong, Weinan Zhang, Fei Fang, Guangjun Zeng, Yue Lin
摘要
As a challenging multi-player card game, DouDizhu has recently drawn much attention for analyzing competition and collaboration in imperfect-information games. In this paper, we propose PerfectDou, a state-of-the-art DouDizhu AI system that dominates the game, in an actor-critic framework with a proposed technique named perfect information distillation. In detail, we adopt a perfecttraining-imperfect-execution framework that allows the agents to utilize the global information to guide the training of the policies as if it is a perfect information game and the trained policies can be used to play the imperfect information game during the actual gameplay. To this end, we characterize card and game features for DouDizhu to represent the perfect and imperfect information. To train our system, we adopt proximal policy optimization with generalized advantage estimation in a parallel training paradigm. In experiments we show how and why PerfectDou beats all existing AI programs, and achieves state-of-the-art performance. * Equal contribution. Yang is responsible for the basic idea, system design and implementation details; Liu mainly contributes to the methodology, writing and experimental design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun 等NeurIPS 2023 · 被引用 24 次
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 被引用 22 次
- Privileged Knowledge State Distillation for Reinforcement Learning-based Educational Path RecommendationQingyao Li, Wei Xia, Li'ang Yin, Jiarui Jin 等KDD 2024 · 被引用 6 次
- Dual Critic Reinforcement Learning under Partial ObservabilityJinqiu Li, Enmin Zhao, Tong Wei, Junliang Xing 等NeurIPS 2024 · 被引用 3 次
- Visual Imitation Learning with Patch RewardsMinghuan Liu, Tairan He, Weinan Zhang, Shuicheng Yan 等ICLR 2023 · 被引用 1 次
它引用的顶会 Paper4
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi 等AAAI 2020 · 被引用 395 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
- Universal Trading for Order Execution with Oracle Policy DistillationYuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou 等AAAI 2021 · 被引用 52 次
- Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information GameHaobo Fu, Weiming Liu, Shuang Wu, Yijia Wang 等ICLR 2022 · 被引用 32 次
相关 Paper
- DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement LearningDaochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang 等ICML 2021 · 被引用 150 次
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
- Joint Policy Search for Multi-agent Collaboration with Imperfect InformationYuandong Tian, Qucheng Gong, Yu JiangNeurIPS 2020 · 被引用 24 次
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray 等ICLR 2023 · 被引用 10 次
- No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained EmbeddingYanchang Fu, Shengda Liu, Pei Xu, Kaiqi HuangAAAI 2026
