AlphaZero-based Proof Cost Network to Aid Game Solving
Ti-Rong Wu, Chung-Chin Shih, Ting-Han Wei, Meng-Yu Tsai, Wei-Yuan Hsu, I-Chen Wu
摘要
The AlphaZero algorithm learns and plays games without hand-crafted expert knowledge. However, since its objective is to play well, we hypothesize that a better objective can be defined for the related but separate task of solving games. This paper proposes a novel approach to solving problems by modifying the training target of the AlphaZero algorithm, such that it prioritizes solving the game quickly, rather than winning. We train a Proof Cost Network (PCN), where proof cost is a heuristic that estimates the amount of work required to solve problems. This matches the general concept of the so-called proof number from proof number search, which has been shown to be well-suited for game solving. We propose two specific training targets. The first finds the shortest path to a solution, while the second estimates the proof cost. We conduct experiments on solving 15x15 Gomoku and 9x9 Killall-Go problems with both MCTS-based and FDFPN solvers. Comparisons between using AlphaZero networks and PCN as heuristics show that PCN can solve more problems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- A Novel Approach to Solving Goal-Achieving Problems for Board GamesChung-Chin Shih, Ti-Rong Wu, Ting-Han Wei, I-Chen WuAAAI 2022 · 被引用 7 次
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert 等ICML 2020 · 被引用 84 次
- Efficient Learning for AlphaZero via Path ConsistencyDengwei Zhao, Shikui Tu, Lei XuICML 2022 · 被引用 8 次
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang 等ICLR 2026
- Policy-Guided Heuristic Search with GuaranteesLaurent Orseau, Levi H. S. LelisAAAI 2021 · 被引用 30 次
