Regret-Guided Search Control for Efficient Learning in AlphaZero
Yun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang, Ti-Rong Wu
Abstract
Reinforcement learning (RL) agents achieve remarkable performance but remain far less learning-efficient than humans. While RL agents require extensive self-play games to extract useful signals, humans often need only a few games, improving rapidly by repeatedly revisiting states where mistakes occurred. This idea, known as search control, aims to restart from valuable states rather than always from the initial state. In AlphaZero, prior work Go-Exploit applies this idea by sampling past states from self-play or search trees, but it treats all states equally, regardless of their learning potential. We propose Regret-Guided Search Control (RGSC), which extends AlphaZero with a regret network that learns to identify high-regret states, where the agent's evaluation diverges most from the actual outcome. These states are collected from both self-play trajectories and MCTS nodes, stored in a prioritized regret buffer, and reused as new starting positions. Across 9x9 Go, 10x10 Othello, and 11x11 Hex, RGSC outperforms AlphaZero and Go-Exploit by an average of 77 and 89 Elo, respectively. When training on a well-trained 9x9 Go model, RGSC further improves the win rate against KataGo from 69.3% to 78.2%, while both baselines show no improvement. These results demonstrate that RGSC provides an effective mechanism for search control, improving both efficiency and robustness of AlphaZero training. Our code is available at https://rlg.iis.sinica.edu.tw/papers/rgsc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41fce70c-e8a6-46e2-b132-32ad7e7c9c48Builds on7
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan et al.ICML 2022 · 175 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain et al.ICML 2021 · 99 citations
Related papers
- Efficient Learning for AlphaZero via Path ConsistencyDengwei Zhao, Shikui Tu, Lei XuICML 2022 · 8 citations
- Uncertainty-Guided Exploration for Efficient AlphaZero TrainingScott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut T. KandemirNeurIPS 2025
- Are AlphaZero-like Agents Robust to Adversarial Perturbations?Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai et al.NeurIPS 2022 · 15 citations
- Evaluation beyond Task Performance: Analyzing Concepts in AlphaZero in HexCharles Lovering, Jessica Zosa Forde, George Konidaris, Ellie Pavlick et al.NeurIPS 2022 · 13 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
