Residual-MPPI: Online Policy Customization for Continuous Control
Pengcheng Wang, Chenran Li, Catherine Weaver, Kenta Kawamoto, Masayoshi Tomizuka, Chen Tang, Wei Zhan
Abstract
Policies developed through Reinforcement Learning (RL) and Imitation Learning (IL) have shown great potential in continuous control tasks, but real-world applications often require adapting trained policies to unforeseen requirements. While fine-tuning can address such needs, it typically requires additional data and access to the original training metrics and parameters. In contrast, an online planning algorithm, if capable of meeting the additional requirements, can eliminate the necessity for extensive training phases and customize the policy without knowledge of the original training scheme or task. In this work, we propose a generic online planning algorithm for customizing continuous-control policies at the execution time, which we call Residual-MPPI. It can customize a given prior policy on new performance metrics in few-shot and even zero-shot online settings, given access to the prior action distribution alone. Through our experiments, we demonstrate that the proposed Residual-MPPI algorithm can accomplish the fewshot/zero-shot online policy customization task effectively, including customizing the champion-level racing agent, Gran Turismo Sophy (GT Sophy) 1.0, in the challenging car racing scenario, Gran Turismo Sport (GTS) environment. Code for MuJoCo experiments is included in the supplementary and will be opensourced upon acceptance. Demo videos and code are available on our website: https://sites.google.com/view/residual-mppi .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- What Makes Value Learning Efficient in Residual Reinforcement Learning?Guozheng Ma, Lu Li, Haoyu Wang, Zixuan Liu et al.ICML 2026 · 2 citations
- DADP: Domain Adaptive Diffusion PolicyPengcheng Wang, Qinghang Liu, Haotian Lin, Yiheng Li et al.ICML 2026 · 1 citation
- A KL-regularization framework for learning to plan with adaptive priorsÁlvaro Serra-Gómez, Daniel Jarne Ornia, Dhruva Tirumala, Thomas M MoerlandICML 2026
Builds on8
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu et al.ICML 2022 · 194 citations
- Deep Imitative Models for Flexible Inference, Planning, and ControlNicholas Rhinehart, Rowan McAllister, Sergey LevineICLR 2020 · 159 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
- CEIL: Generalized Contextual Imitation LearningJinxin Liu, Li He, Yachen Kang, Zifeng Zhuang et al.NeurIPS 2023 · 23 citations
Related papers
- Residual Q-Learning: Offline and Online Policy Customization without ValueChenran Li, Chen Tang, Haruki Nishimura, Jean Mercat et al.NeurIPS 2023 · 15 citations
- Visual Reinforcement Learning with Residual ActionZhenxian Liu, Peixi Peng, Yonghong TianAAAI 2025 · 4 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and PlanningYizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang et al.ICML 2024 · 5 citations
- OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesSarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena et al.ICML 2026
