Convex Regularization in Monte-Carlo Tree Search
Tuan Dam, Carlo D'Eramo, Jan Peters, Joni Pajarinen
Abstract
Monte-Carlo planning and Reinforcement Learning (RL) are essential to sequential decision making. The recent AlphaGo and AlphaZero algorithms have shown how to successfully combine these two paradigms to solve large scale sequential decision problems. These methodologies exploit a variant of the well-known UCT algorithm to trade off the exploitation of good actions and the exploration of unvisited states, but their empirical success comes at the cost of poor sampleefficiency and high computation time. In this paper, we overcome these limitations by introducing the use of convex regularization in Monte-Carlo Tree Search (MCTS) to drive exploration efficiently and to improve policy updates. First, we introduce a unifying theory on the use of generic convex regularizers in MCTS, deriving the first regret analysis of regularized MCTS and showing that it guarantees an exponential convergence rate. Second, we exploit our theoretical framework to introduce novel regularized backup operators for MCTS, based on the relative entropy of the policy update and, more importantly, on the Tsallis entropy of the policy, for which we prove superior theoretical guarantees. We empirically verify the consequence of our theoretical results on a toy problem. Finally, we show how our framework can easily be incorporated in AlphaGo and we empirically show the superiority of convex regularization, w.r.t. representative baselines, on wellknown RL problems across several Atari games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9834c757-49e0-4e96-993c-382711e7f3e4Cited by top-tier papers5
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Monte Carlo Tree Search with Boltzmann ExplorationMichael Painter, Mohamed Baioumy, Nick Hawes, Bruno LacerdaNeurIPS 2023 · 17 citations
- KeeA*: Epistemic Exploratory A* Search via Knowledge CalibrationDengwei Zhao, Shikui Tu, Yanan Sun, Lei XuNeurIPS 2025
- Online Robust Reinforcement Learning Through Monte-Carlo PlanningTuan Dam, Kishan Panaganti, Brahim Driss, Adam WiermanICML 2025
- Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic EnvironmentsKhang Luong, Nam Nguyen, Hoang Ta, Hung Tran-The et al.ICML 2026
Builds on1
Related papers
- Provably Efficient Long-Horizon Exploration in Monte Carlo Tree Search through State Occupancy RegularizationLiam Schramm, Abdeslam BoulariasICML 2024 · 1 citation
- Recursive Monte-Carlo Tree SearchBenjamin Howard, Keith FrankstonICML 2026
- Accelerating Monte Carlo Tree Search with Probability Tree State AbstractionYangqing Fu, Ming Sun, Buqing Nie, Yue GaoNeurIPS 2023 · 5 citations
- Goal-Directed Planning via Hindsight Experience ReplayLorenzo Moro, Amarildo Likmeta, Enrico Prati, Marcello RestelliICLR 2022 · 14 citations
- POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic AnalysisWeichao Mao, Kaiqing Zhang, Qiaomin Xie, Tamer BasarNeurIPS 2020 · 18 citations
