Accelerating and Improving AlphaZero Using Population Based Training
Ti-Rong Wu, Ting-Han Wei, I-Chen Wu
Abstract
AlphaZero has been very successful in many games. Unfortunately, it still consumes a huge amount of computing resources, the majority of which is spent in self-play. Hyperparameter tuning exacerbates the training cost since each hyperparameter configuration requires its own time to train one run, during which it will generate its own self-play records. As a result, multiple runs are usually needed for different hyperparameter configurations. This paper proposes using population based training (PBT) to help tune hyperparameters dynamically and improve strength during training time. Another significant advantage is that this method requires a single run only, while incurring a small additional time cost, since the time for generating self-play records remains unchanged though the time for optimization is increased following the AlphaZero training algorithm. In our experiments for 9x9 Go, the PBT method is able to achieve a higher win rate for 9x9 Go than the baselines, each with its own hyperparameter configuration and trained individually. For 19x19 Go, with PBT, we are able to obtain improvements in playing strength. Specifically, the PBT agent can obtain up to 74% win rate against ELF OpenGo, an open-source state-of-the-art AlphaZero program using a neural network of a comparable capacity. This is compared to a saturated non-PBT agent, which achieves a win rate of 47% against ELF OpenGo under the same circumstances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c3d3a6c-443e-4762-b678-e83020005deeCited by top-tier papers3
- Are AlphaZero-like Agents Robust to Adversarial Perturbations?Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai et al.NeurIPS 2022 · 15 citations
- A Novel Approach to Solving Goal-Achieving Problems for Board GamesChung-Chin Shih, Ti-Rong Wu, Ting-Han Wei, I-Chen WuAAAI 2022 · 7 citations
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player GamesKazuki Ota, Takayuki Osa, Motoki Omura, Tatsuya HaradaICML 2026
Related papers
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 105 citations
- Iterated Population Based Training with Task-Agnostic RestartsAlexander Chebykin, Tanja Alderliesten, Peter A.N BosmanICML 2026
- Enhancing Chess Reinforcement Learning with Graph RepresentationTomas Rigaux, Hisashi KashimaNeurIPS 2024 · 5 citations
- Multi-Objective Population Based TrainingArkadiy Dushatskiy, Alexander Chebykin, Tanja Alderliesten, Peter A. N. BosmanICML 2023 · 4 citations
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 3 citations
