Stop! Planner Time: Metareasoning for Probabilistic Planning Using Learned Performance Profiles
Matthew Budd, Bruno Lacerda, Nick Hawes
摘要
The metareasoning framework aims to enable autonomous agents to factor in planning costs when making decisions. In this work, we develop the first non-myopic metareasoning algorithm for planning with Markov decision processes. Our method learns the behaviour of anytime probabilistic planning algorithms from performance data. Specifically, we propose a novel model for metareasoning, based on contextual performance profiles that predict the value of the planner's current solution given the time spent planning, the state of the planning algorithm's internal parameters, and the difficulty of the planning problem being solved. This model removes the need to assume that the current solution quality is always known, broadening the class of metareasoning problems that can be addressed. We then employ deep reinforcement learning to learn a policy that decides, at each timestep, whether to continue planning or start executing the current plan, and how to set hyperparameters of the planner to enhance its performance. We demonstrate our algorithm's ability to perform effective metareasoning in two domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 被引用 4 次
- Improving planning and MBRL with temporally-extended actionsPalash Chatterjee, Roni KhardonNeurIPS 2025
- From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in LRMs via Decoupled Reasoning and ControlRui Ha, Rui Pu, Chaozhuo Li, Li Sun 等ACL 2026 · 被引用 5 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- People Do Not Just Plan, They Plan to PlanMark K. Ho, David Abel, Jonathan D. Cohen, Michael L. Littman 等AAAI 2020 · 被引用 1 次
