Scalable Safe Policy Improvement via Monte Carlo Tree Search
Alberto Castellini, Federico Bianchi, Edoardo Zorzi, Thiago D. Simão, Alessandro Farinelli, Matthijs T. J. Spaan
摘要
Algorithms for safely improving policies are important to deploy reinforcement learning approaches in real-world scenarios. In this work, we propose an algorithm, called MCTS-SPIBB, that computes safe policy improvement online using a Monte Carlo Tree Search based strategy. We theoretically prove that the policy generated by MCTS-SPIBB converges, as the number of simulations grows, to the optimal safely improved policy generated by Safe Policy Improvement with Baseline Bootstrapping (SPIBB), a popular algorithm based on policy iteration. Moreover, our empirical analysis performed on three standard benchmark domains shows that MCTS-SPIBB scales to significantly larger problems than SPIBB because it computes the policy online and locally, i.e., only in the states actually visited by the agent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deep SPI: Safe Policy Improvement via World ModelsFlorent Delgrange, Raphaël Avalos, Willem RöpkeICLR 2026 · 被引用 4 次
- Scalable Safe Policy Improvement for Factored Multi-Agent MDPsFederico Bianchi, Edoardo Zorzi, Alberto Castellini, Thiago D. Simão 等ICML 2024 · 被引用 3 次
- Weak-to-Strong Generalization with Failure TrajectoriesRuimeng Ye, Zihan Wang, Yang Xiao, Zinan Ling 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper4
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Towards Safe Policy Improvement for Non-Stationary MDPsYash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White 等NeurIPS 2020 · 被引用 47 次
相关 Paper
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPsHarsh Satija, Philip S. Thomas, Joelle Pineau, Romain LarocheNeurIPS 2021 · 被引用 30 次
- CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold PoliciesBrian M. Cho, Ana-Roxana Pop, Kyra Gan, Sam Corbett-Davies 等KDD 2025 · 被引用 2 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Safe Offline Reinforcement Learning with Real-Time Budget ConstraintsQian Lin, Bo Tang, Zifan Wu, Chao Yu 等ICML 2023 · 被引用 31 次
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 被引用 78 次
