Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
Adrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu, Hao Su
Abstract
Long-horizon tasks in robotic manipulation present significant challenges in reinforcement learning (RL) due to the difficulty of designing dense reward functions and effectively exploring the expansive state-action space. However, despite a lack of dense rewards, these tasks often have a multi-stage structure, which can be leveraged to decompose the overall objective into manageable subgoals. In this work, we propose Demonstration-Augmented Reward, Policy, and World Model Learning (DEMO 3 ), a framework that exploits this structure for efficient learning from visual inputs. Specifically, our approach incorporates multi-stage dense reward learning, a bi-phasic training scheme, and world model learning into a carefully designed demonstrationaugmented RL framework that strongly mitigates the challenge of exploration in long-horizon tasks. Our evaluations demonstrate that our method improves data-efficiency by an average of 40% and by 70% on particularly difficult tasks compared to state-of-the-art approaches. We validate this across 16 sparse-reward tasks spanning four domains, including challenging humanoid visual control tasks using as few as five demonstrations. Website with code and visualizations can be found at https://adrialopezescoriza.github.io/demo3 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ff6c219-ab55-4224-93ba-a8de2e71ec1aCited by top-tier papers1
Ask how each one uses itBuilds on11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
Related papers
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 72 citations
- MoDem: Accelerating Visual Model-Based Reinforcement Learning with DemonstrationsNicklas Hansen, Yixin Lin, Hao Su, Xiaolong Wang et al.ICLR 2023 · 8 citations
- DrS: Learning Reusable Dense Rewards for Multi-Stage TasksTongzhou Mu, Minghua Liu, Hao SuICLR 2024 · 8 citations
- Skill-based Meta-Reinforcement LearningTaewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang et al.ICLR 2022 · 55 citations
- DemoFunGrasp: Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement LearningChuan Mao, Haoqi Yuan, Ziye Huang, Chaoyi Xu et al.CVPR 2026
