DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
Tongzhou Mu, Minghua Liu, Hao Su
Abstract
The success of many RL techniques heavily relies on human-engineered dense rewards, which typically demand substantial domain expertise and extensive trial and error. In our work, we propose DrS (Dense reward learning from Stages), a novel approach for learning reusable dense rewards for multi-stage tasks in a data-driven manner. By leveraging the stage structures of the task, DrS learns a high-quality dense reward from sparse rewards and demonstrations if given. The learned rewards can be reused in unseen tasks, thus reducing the human effort for reward engineering. Extensive experiments on three physical robot manipulation task families with 1000+ task variants demonstrate that our learned rewards can be reused in unseen tasks, resulting in improved performance and sample efficiency of RL algorithms. The learned rewards even achieve comparable performance to human-engineered rewards on some tasks. See our project page (https://sites.google.com/view/iclr24drs) for more details.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d17dc385-b134-4ef0-aa7b-56ac609ec7cbCited by top-tier papers5
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsSenyu Fei, Siyin Wang, Li Ji, Ao Li et al.CVPR 2026 · 28 citations
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon TasksMingxuan Yan, Yuping Wang, Zechun Liu, Jiachen LiNeurIPS 2025 · 4 citations
- Cooperative Bargaining Games Without Utilities: Mediated Solutions from Direction OraclesKushagra Gupta, Surya Murthy, Mustafa O. Karabag, Ufuk Topcu et al.NeurIPS 2025 · 2 citations
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL ProblemsEgor Cherepanov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 1 citation
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee et al.ICLR 2025
Builds on5
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent et al.NeurIPS 2021 · 106 citations
- State Alignment-based Imitation LearningFangchen Liu, Zhan Ling, Tongzhou Mu, Hao SuICLR 2020 · 103 citations
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling et al.ICLR 2023 · 21 citations
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 10 citations
- Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard et al.CVPR 2020
Related papers
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu et al.ICML 2025
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 72 citations
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu et al.ICLR 2024 · 142 citations
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal DistanceYuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman et al.ICML 2026 · 7 citations
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen et al.NeurIPS 2025 · 3 citations
