Uncertainty-Guided Exploration for Efficient AlphaZero Training
Scott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. Kandemir
摘要
AlphaZero has achieved remarkable success in complex decision-making problems through self-play and neural network training. However, its self-play process remains inefficient due to limited exploration of high-uncertainty positions, the overlooked runner-up decisions in Monte Carlo Tree Search (MCTS), and high variance in value labels. To address these challenges, we propose and evaluate uncertainty-guided exploration by branching from high-uncertainty positions using our proposed Label Change Rate (LCR) metric, which is further refined by a Bayesian inference framework. Our proposed approach leverages runner-up MCTS decisions to create multiple variations, and ensembles value labels across these variations to reduce variance. We investigate three key design parameters for our branching strategy: where to branch, how many variations to branch, and which move to play in the new branch. Our empirical findings indicate that branching with 10 variations per game provides the best performance-exploration balance. Overall, our end-to-end results show an improved sample efficiency over the baseline by 58.5% on 9x9 Go in the early stage of training and by 47.3% on 19x19 Go in the late stage of training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel 等NeurIPS 2021 · 被引用 345 次
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 被引用 84 次
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 被引用 53 次
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataShengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You 等ICML 2024 · 被引用 36 次
相关 Paper
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang 等ICLR 2026
- Learning to Stop: Dynamic Simulation Monte-Carlo Tree SearchLi-Cheng Lan, Ti-Rong Wu, I-Chen Wu, Cho-Jui HsiehAAAI 2021 · 被引用 7 次
- Speculative Monte-Carlo Tree SearchScott Cheng, Mahmut T. Kandemir, Ding-Yong HongNeurIPS 2024 · 被引用 4 次
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
- Are AlphaZero-like Agents Robust to Adversarial Perturbations?Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai 等NeurIPS 2022 · 被引用 15 次
