Bootstrapped Model Predictive Control
Yuhang Wang, Hanwei Guo, Sizhe Wang, Long Qian, Xuguang Lan
摘要
Model Predictive Control (MPC) has been demonstrated to be effective in continuous control tasks. When a world model and a value function are available, planning a sequence of actions ahead of time leads to a better policy. Existing methods typically obtain the value function and the corresponding policy in a model-free manner. However, we find that such an approach struggles with complex tasks, resulting in poor policy learning and inaccurate value estimation. To address this problem, we leverage the strengths of MPC itself. In this work, we introduce Bootstrapped Model Predictive Control (BMPC), a novel algorithm that performs policy learning in a bootstrapped manner. BMPC learns a network policy by imitating an MPC expert, and in turn, uses this policy to guide the MPC process. Combined with model-based TD-learning, our policy learning yields better value estimation and further boosts the efficiency of MPC. We also introduce a lazy reanalyze mechanism, which enables computationally efficient imitation learning. Our method achieves superior performance over prior works on diverse continuous control tasks. In particular, on challenging high-dimensional locomotion tasks, BMPC significantly improves data efficiency while also enhancing asymptotic performance and training stability, with comparable training time and smaller network sizes. Code is available at https://github.com/wertyuilife2/bmpc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj 等NeurIPS 2025 · 被引用 27 次
- Bootstrap Off-policy with World ModelGuojian Zhan, Likun Wang, Xiangteng Zhang, Jiaxin Gao 等NeurIPS 2025 · 被引用 9 次
- Langevin Rollout Optimization for Modelic Reinforcement LearningTianyi Zhang, Likun Wang, Guojian Zhan, Feihong Zhang 等ICML 2026 · 被引用 7 次
- The Surprising Difficulty of Search in Model-Based Reinforcement LearningWei-Di Chang, Mikael Henaff, Brandon Amos, Gregory Dudek 等ICML 2026 · 被引用 4 次
- Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationJonathan Spieler, Sven BehnkeICML 2026
它引用的顶会 Paper6
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited DataShengjie Wang, Shaohuai Liu, Weirui Ye, Jiacheng You 等ICML 2024 · 被引用 36 次
- Transformer-based World Models Are Happy With 100k InteractionsJan Robine, Marc Höftmann, Tobias Uelwer, Stefan HarmelingICLR 2023 · 被引用 4 次
相关 Paper
- Evaluating Model-Based Planning and Planner Amortization for Continuous ControlArunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza 等ICLR 2022 · 被引用 18 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- On the Expressivity of Neural Networks for Deep Reinforcement LearningKefan Dong, Yuping Luo, Tianhe Yu, Chelsea Finn 等ICML 2020 · 被引用 33 次
- Blending MPC & Value Function Approximation for Efficient Reinforcement LearningMohak Bhardwaj, Sanjiban Choudhury, Byron BootsICLR 2021 · 被引用 3 次
- Bisimulation Metric for Model Predictive ControlYutaka Shimizu, Masayoshi TomizukaICLR 2025
