Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Xiu Yuan, Tongzhou Mu, Stone Tao, Yunhao Fang, Mengke Zhang, Hao Su
摘要
Recent advancements in robot learning have used imitation learning with large models and extensive demonstrations to develop effective policies. However, these models are often limited by the quantity, quality, and diversity of demonstrations. This paper explores improving offline-trained imitation learning models through online interactions with the environment. We introduce Policy Decorator, which uses a model-agnostic residual policy to refine large imitation learning models during online interactions. By implementing controlled exploration strategies, Policy Decorator enables stable, sample-efficient online learning. Our evaluation spans eight tasks across two benchmarks-ManiSkill and Adroit-and involves two state-of-the-art imitation learning models (Behavior Transformer and Diffusion Policy). The results show Policy Decorator effectively improves the offline-trained policies and preserves the smooth motion of imitation learning models, avoiding the erratic behaviors of pure RL policies. See our project page for videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human CorrectionsXiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran SongNeurIPS 2025 · 被引用 57 次
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
- EXPO: Stable Reinforcement Learning with Expressive PoliciesPerry Dong, Qiyang Li, Dorsa Sadigh, Chelsea FinnICLR 2026 · 被引用 35 次
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj 等NeurIPS 2025 · 被引用 27 次
- RFS: Reinforcement learning with Residual flow steering for dexterous manipulationEntong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek GuptaICLR 2026 · 被引用 13 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
相关 Paper
- Curriculum Offline Imitating LearningMinghuan Liu, Hanye Zhao, Zhengyu Yang, Jian Shen 等NeurIPS 2021 · 被引用 5 次
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal TransportMingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin WangICML 2025
- Translating Flow to Policy via Hindsight Online ImitationYitian Zheng, Zhangchen Ye, Weijun Dong, Shengjie Wang 等ICLR 2026 · 被引用 2 次
- Iterative Regularized Policy Optimization with Imperfect DemonstrationsXudong Gong, Dawei Feng, Kele Xu, Yuanzhao Zhai 等ICML 2024 · 被引用 5 次
- Residual Q-Learning: Offline and Online Policy Customization without ValueChenran Li, Chen Tang, Haruki Nishimura, Jean Mercat 等NeurIPS 2023 · 被引用 15 次
