Power Mean Estimation in Stochastic Continuous Monte-Carlo Tree Search
Tuan Dam
Abstract
We consider Monte-Carlo Tree Search (MCTS) applied to Markov Decision Processes (MDPs) and Partially Observable MDPs (POMDPs), and the well-known Upper Confidence bound for Trees (UCT) algorithm. In UCT, a tree with nodes (states) and edges (actions) is incrementally built by the expansion of nodes, and the values of nodes are updated through a backup strategy based on the average value of child nodes. However, it has been shown that with enough samples the maximum operator yields more accurate node value estimates than averaging. Instead of settling for one of these value estimates, we go a step further proposing a novel backup strategy which uses the power mean operator, which computes a value between the average and maximum value. We call our new approach Power-UCT, and argue how the use of the power mean operator helps to speed up the learning in MCTS. We theoretically analyze our method providing guarantees of convergence to the optimum. Finally, we empirically demonstrate the effectiveness of our method in well-known MDP and POMDP benchmarks, showing significant improvement in performance and convergence speed w.r.t. state of the art algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cd43b37-e555-4bd8-8f45-65d8f5132cbcCited by top-tier papers1
Ask how each one uses itRelated papers
- Monte-Carlo Tree Search with Uncertainty Propagation via Optimal TransportTuan Dam, Pascal Stenger, Lukas Schneider, Joni Pajarinen et al.ICML 2025
- Threshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto CurvesMartin Kurecka, Václav Nevyhostený, Petr Novotný, Vít UncovskýAAAI 2025 · 1 citation
- Monte Carlo Tree Search with Boltzmann ExplorationMichael Painter, Mohamed Baioumy, Nick Hawes, Bruno LacerdaNeurIPS 2023 · 17 citations
- Online Robust Reinforcement Learning Through Monte-Carlo PlanningTuan Dam, Kishan Panaganti, Brahim Driss, Adam WiermanICML 2025
- Online POMDP Planning with Anytime Deterministic GuaranteesMoran Barenboim, Vadim IndelmanNeurIPS 2023 · 13 citations
