Uncertainty-Guided Exploration for Efficient AlphaZero Training
Scott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. Kandemir
Abstract
AlphaZero has achieved remarkable success in complex decision-making problems through self-play and neural network training. However, its self-play process remains inefficient due to limited exploration of high-uncertainty positions, the overlooked runner-up decisions in Monte Carlo Tree Search (MCTS), and high variance in value labels. To address these challenges, we propose and evaluate uncertainty-guided exploration by branching from high-uncertainty positions using our proposed Label Change Rate (LCR) metric, which is further refined by a Bayesian inference framework. Our proposed approach leverages runner-up MCTS decisions to create multiple variations, and ensembles value labels across these variations to reduce variance. We investigate three key design parameters for our branching strategy: where to branch, how many variations to branch, and which move to play in the new branch. Our empirical findings indicate that branching with 10 variations per game provides the best performance-exploration balance. Overall, our end-to-end results show an improved sample efficiency over the baseline by 58.5% on 9x9 Go in the early stage of training and by 47.3% on 19x19 Go in the late stage of training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 168812ae-ca8e-490b-b744-da9c55f00489Builds on11
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 53 citations
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataShengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You et al.ICML 2024 · 36 citations
Related papers
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang et al.ICLR 2026
- Learning to Stop: Dynamic Simulation Monte-Carlo Tree SearchLi-Cheng Lan, Ti-Rong Wu, I-Chen Wu, Cho-Jui HsiehAAAI 2021 · 7 citations
- Speculative Monte-Carlo Tree SearchScott Cheng, Mahmut T. Kandemir, Ding-Yong HongNeurIPS 2024 · 4 citations
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
- Are AlphaZero-like Agents Robust to Adversarial Perturbations?Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai et al.NeurIPS 2022 · 15 citations
