Efficient Learning for AlphaZero via Path Consistency
Dengwei Zhao, Shikui Tu, Lei Xu
摘要
In recent years, deep reinforcement learning have made great breakthroughs on board games. Still, most of the works require huge computational resources for a large scale of environmental interactions or self-play for the games. This paper aims at building powerful models under a limited amount of self-plays which can be utilized by a human throughout the lifetime. We proposes a learning algorithm built on AlphaZero, with its path searching regularised by a path consistency (PC) optimality, i.e., values on one optimal search path should be identical. Thus, the algorithm is shortly named PCZero. In implementation, historical trajectory and scouted search paths by MCTS makes a good balance between exploration and exploitation, which enhances the generalization ability effectively. PCZero obtains 94.1% winning rate against the champion of Hex Computer Olympiad in 2015 on 13 × 13 Hex, much higher than 84.3% by AlphaZero. The models consume only 900K self-play games, about the amount humans can study in a lifetime. The improvements by PCZero have been also generalized to Othello and Gomoku. Experiments also demonstrate the efficiency of PCZero under offline learning setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generalized Weighted Path Consistency for Mastering Atari GamesDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2023 · 被引用 5 次
- SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2024 · 被引用 4 次
- Uncertainty-Guided Exploration for Efficient AlphaZero TrainingScott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. KandemirNeurIPS 2025
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player GamesKazuki Ota, Takayuki Osa, Motoki Omura, Tatsuya HaradaICML 2026
相关 Paper
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang 等ICLR 2026
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 被引用 3 次
- Deep Reinforcement Learning for General Game PlayingAdrian Goldwaser, Michael ThielscherAAAI 2020 · 被引用 46 次
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 被引用 84 次
- Evaluation beyond Task Performance: Analyzing Concepts in AlphaZero in HexCharles Lovering, Jessica Zosa Forde, George Konidaris, Ellie Pavlick 等NeurIPS 2022 · 被引用 13 次
