Threshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto Curves
Martin Kurecka, Václav Nevyhostený, Petr Novotný, Vít Uncovský
摘要
Constrained Markov decision processes (CMDPs), in which the agent optimizes expected payoffs while keeping the expected cost below a given threshold, are the leading framework for safe sequential decision making under stochastic uncertainty. Among algorithms for planning and learning in CMDPs, methods based on Monte Carlo tree search (MCTS) have particular importance due to their efficiency and extendibility to more complex frameworks (such as partially observable settings and games). However, current MCTS-based methods for CMDPs either struggle with finding safe (i.e., constraint-satisfying) policies, or are too conservative and do not find valuable policies. We introduce Threshold UCT (T-UCT), an online MCTS-based algorithm for CMDP planning. Unlike previous MCTS-based CMDP planners, T-UCT explicitly estimates Pareto curves of cost-utility trade-offs throughout the search tree, using these together with a novel action selection and threshold update rule to seek safe and valuable policies. Our experiments demonstrate that our approach significantly outperforms state-of-the-art methods from the literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Qualitative Controller Synthesis for Consumption Markov Decision ProcessesFrantisek Blahoudek, Tomás Brázdil, Petr Novotný, Melkior Ornik 等CAV 2020 · 被引用 9 次
- Reinforcement Learning of Risk-Constrained Policies in Markov Decision ProcessesTomás Brázdil, Krishnendu Chatterjee, Petr Novotný, Jiri VahalaAAAI 2020 · 被引用 5 次
- Fuel in Markov Decision Processes (FiMDP): A Practical Approach to ConsumptionFrantisek Blahoudek, Murat Cubuktepe, Petr Novotný, Melkior Ornik 等FM 2021 · 被引用 2 次
相关 Paper
- A Sample-Efficient Algorithm for Episodic Finite-Horizon MDP with ConstraintsKrishna Chaitanya Kalagarla, Rahul Jain, Pierluigi NuzzoAAAI 2021 · 被引用 58 次
- Power Mean Estimation in Stochastic Continuous Monte-Carlo Tree SearchTuan DamICML 2025
- Monte-Carlo Tree Search in Continuous Action Spaces with Value GradientsJongmin Lee, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung KimAAAI 2020 · 被引用 24 次
- Upper Confidence Primal-Dual Reinforcement Learning for CMDP with Adversarial LossShuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye 等NeurIPS 2020 · 被引用 65 次
- Online Learning in CMDPs: Handling Stochastic and Adversarial ConstraintsFrancesco Emanuele Stradi, Jacopo Germano, Gianmarco Genalti, Matteo Castiglioni 等ICML 2024 · 被引用 7 次
