Online Robust Reinforcement Learning Through Monte-Carlo Planning
Tuan Dam, Kishan Panaganti, Brahim Driss, Adam Wierman
Abstract
Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decisionmaking problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the real-world scenarios incur ambiguity due to their modeling mismatches in low-fidelity simulators. In this work, we present a new robust variant of MCTS that mitigates dynamical model ambiguities. Our algorithm addresses transition dynamics and reward distribution ambiguities to bridge the gap between simulation-based planning and real-world deployment. We incorporate a robust power mean backup operator and carefully designed exploration bonuses to ensure finite-sample convergence at every node in the search tree. We show that our algorithm achieves a convergence rate of O(n -1/2 ) for the value estimation at the root node, comparable to that of standard MCTS. Finally, we provide empirical evidence that our method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94c18e25-7b02-43c9-bed1-060ca9381df3Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 78 citations
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong et al.ICML 2022 · 72 citations
- Policy Gradient for Rectangular Robust Markov Decision ProcessesNavdeep Kumar, Esther Derman, Matthieu Geist, Kfir Y. Levy et al.NeurIPS 2023 · 45 citations
- POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic AnalysisWeichao Mao, Kaiqing Zhang, Qiaomin Xie, Tamer BasarNeurIPS 2020 · 18 citations
Related papers
- Monte Carlo Tree Search in the Presence of Transition UncertaintyFarnaz Kohankhaki, Kiarash Aghakasiri, Hongming Zhang, Ting-Han Wei et al.AAAI 2024 · 4 citations
- Power Mean Estimation in Stochastic Continuous Monte-Carlo Tree SearchTuan DamICML 2025
- Monte-Carlo Tree Search with Uncertainty Propagation via Optimal TransportTuan Dam, Pascal Stenger, Lukas Schneider, Joni Pajarinen et al.ICML 2025
- Extreme Value Monte Carlo Tree Search for Classical PlanningMasataro Asai, Stephen WissowAAAI 2026 · 2 citations
- Online Robust Planning Under Model Uncertainty: A Sample-Based ApproachTamir Shazman, Idan Lev-Yehudi, Ron Benchetrit, Vadim IndelmanAAAI 2026
