Online Robust Reinforcement Learning Through Monte-Carlo Planning
Tuan Dam, Kishan Panaganti, Brahim Driss, Adam Wierman
摘要
Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decisionmaking problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the real-world scenarios incur ambiguity due to their modeling mismatches in low-fidelity simulators. In this work, we present a new robust variant of MCTS that mitigates dynamical model ambiguities. Our algorithm addresses transition dynamics and reward distribution ambiguities to bridge the gap between simulation-based planning and real-world deployment. We incorporate a robust power mean backup operator and carefully designed exploration bonuses to ensure finite-sample convergence at every node in the search tree. We show that our algorithm achieves a convergence rate of O(n -1/2 ) for the value estimation at the root node, comparable to that of standard MCTS. Finally, we provide empirical evidence that our method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 被引用 130 次
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 被引用 78 次
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong 等ICML 2022 · 被引用 72 次
- Policy Gradient for Rectangular Robust Markov Decision ProcessesNavdeep Kumar, Esther Derman, Matthieu Geist, Kfir Y. Levy 等NeurIPS 2023 · 被引用 45 次
- POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic AnalysisWeichao Mao, Kaiqing Zhang, Qiaomin Xie, Tamer BasarNeurIPS 2020 · 被引用 18 次
相关 Paper
- Monte Carlo Tree Search in the Presence of Transition UncertaintyFarnaz Kohankhaki, Kiarash Aghakasiri, Hongming Zhang, Ting-Han Wei 等AAAI 2024 · 被引用 4 次
- Power Mean Estimation in Stochastic Continuous Monte-Carlo Tree SearchTuan DamICML 2025
- Monte-Carlo Tree Search with Uncertainty Propagation via Optimal TransportTuan Dam, Pascal Stenger, Lukas Schneider, Joni Pajarinen 等ICML 2025
- Extreme Value Monte Carlo Tree Search for Classical PlanningMasataro Asai, Stephen WissowAAAI 2026 · 被引用 2 次
- Online Robust Planning Under Model Uncertainty: A Sample-Based ApproachTamir Shazman, Idan Lev-Yehudi, Ron Benchetrit, Vadim IndelmanAAAI 2026
