Efficient Learning for AlphaZero via Path Consistency
Dengwei Zhao, Shikui Tu, Lei Xu
Abstract
In recent years, deep reinforcement learning have made great breakthroughs on board games. Still, most of the works require huge computational resources for a large scale of environmental interactions or self-play for the games. This paper aims at building powerful models under a limited amount of self-plays which can be utilized by a human throughout the lifetime. We proposes a learning algorithm built on AlphaZero, with its path searching regularised by a path consistency (PC) optimality, i.e., values on one optimal search path should be identical. Thus, the algorithm is shortly named PCZero. In implementation, historical trajectory and scouted search paths by MCTS makes a good balance between exploration and exploitation, which enhances the generalization ability effectively. PCZero obtains 94.1% winning rate against the champion of Hex Computer Olympiad in 2015 on 13 × 13 Hex, much higher than 84.3% by AlphaZero. The models consume only 900K self-play games, about the amount humans can study in a lifetime. The improvements by PCZero have been also generalized to Othello and Gomoku. Experiments also demonstrate the efficiency of PCZero under offline learning setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5520a9b9-43af-43ed-89cb-9f63c5772399Cited by top-tier papers4
- Generalized Weighted Path Consistency for Mastering Atari GamesDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2023 · 5 citations
- SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2024 · 4 citations
- Uncertainty-Guided Exploration for Efficient AlphaZero TrainingScott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. KandemirNeurIPS 2025
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player GamesKazuki Ota, Takayuki Osa, Motoki Omura, Tatsuya HaradaICML 2026
Related papers
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang et al.ICLR 2026
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 3 citations
- Deep Reinforcement Learning for General Game PlayingAdrian Goldwaser, Michael ThielscherAAAI 2020 · 46 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Evaluation beyond Task Performance: Analyzing Concepts in AlphaZero in HexCharles Lovering, Jessica Zosa Forde, George Konidaris, Ellie Pavlick et al.NeurIPS 2022 · 13 citations
