Trust-Region Twisted Policy Improvement
Joery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. Spaan
摘要
Monte-Carlo tree search (MCTS) has driven many recent breakthroughs in deep reinforcement learning (RL). However, scaling MCTS to parallel compute has proven challenging in practice which has motivated alternative planners like sequential Monte-Carlo (SMC). Many of these SMC methods adopt particle filters for smoothing through a reformulation of RL as a policy inference problem. Yet, persisting design choices of these particle filters often conflict with the aim of online planning in RL, which is to obtain a policy improvement at the start of planning. Drawing inspiration from MCTS, we tailor SMC planners specifically to RL by improving data generation within the planner through constrained action sampling and explicit terminal state handling, as well as improving policy and value target estimation. This leads to our Trust-Region Twisted SMC (TRT-SMC), which shows improved runtime and sample-efficiency over baseline MCTS and SMC methods in both discrete and continuous domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mitigating Premature Exploitation in Particle-based Monte Carlo for Inference-Time ScalingGiorgio Giannone, Guangxuan Xu, Nikhil Nayak, Rohan Awhad 等ICML 2026 · 被引用 2 次
- Twice Sequential Monte Carlo for Tree SearchYaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper13
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel 等NeurIPS 2021 · 被引用 345 次
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain 等ICML 2021 · 被引用 99 次
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 被引用 84 次
相关 Paper
- Monte-Carlo Tree Search in Continuous Action Spaces with Value GradientsJongmin Lee, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung KimAAAI 2020 · 被引用 24 次
- SPO: Sequential Monte Carlo Policy OptimisationMatthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth 等NeurIPS 2024 · 被引用 8 次
- Monte Carlo Tree Diffusion for System 2 PlanningJaesik Yoon, Hyeonseo Cho, Doojin Baek, Yoshua Bengio 等ICML 2025
- Scalable Safe Policy Improvement via Monte Carlo Tree SearchAlberto Castellini, Federico Bianchi, Edoardo Zorzi, Thiago D. Simão 等ICML 2023 · 被引用 9 次
- Fast Monte Carlo Tree Diffusion: 100× Speedup via Parallel and Sparse PlanningJaesik Yoon, Hyeonseo Cho, Yoshua Bengio, Sungjin AhnNeurIPS 2025
