Monte-Carlo Tree Search in Continuous Action Spaces with Value Gradients
Jongmin Lee, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung Kim
摘要
Monte-Carlo Tree Search (MCTS) is the state-of-the-art online planning algorithm for large problems with discrete action spaces. However, many real-world problems involve continuous action spaces, where MCTS is not as effective as in discrete action spaces. This is mainly due to common practices such as coarse discretization of the entire action space and failure to exploit local smoothness. In this paper, we introduce Value-Gradient UCT (VG-UCT), which combines traditional MCTS with gradient-based optimization of action particles. VG-UCT simultaneously performs a global search via UCT with respect to the finitely sampled set of actions and performs a local improvement via action value gradients. In the experiments, we demonstrate that our approach outperforms existing MCTS methods and other strong baseline algorithms for continuous action spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Bayesian Optimized Monte Carlo PlanningJohn Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji 等AAAI 2021 · 被引用 33 次
- Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Wen-Tse Chen, Jeff SchneiderICLR 2026 · 被引用 10 次
- Monte Carlo Planning with Large Language Model for Text-Based Game AgentsZijing Shi, Meng Fang, Ling ChenICLR 2025
- Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic DesignZhi Zheng, Zhuoliang Xie, Zhenkun Wang, Bryan HooiICML 2025
相关 Paper
- Monte Carlo Tree Search in Continuous Spaces Using Voronoi Optimistic Optimization with Regret BoundsBeomjoon Kim, Kyungjae Lee, Sungbin Lim, Leslie Pack Kaelbling 等AAAI 2020 · 被引用 55 次
- Threshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto CurvesMartin Kurecka, Václav Nevyhostený, Petr Novotný, Vít UncovskýAAAI 2025 · 被引用 1 次
- Trust-Region Twisted Policy ImprovementJoery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. SpaanICML 2025
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 被引用 42 次
- Watch the Unobserved: A Simple Approach to Parallelizing Monte Carlo Tree SearchAnji Liu, Jianshu Chen, Mingze Yu, Yu Zhai 等ICLR 2020 · 被引用 39 次
