DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
Tongzhou Mu, Minghua Liu, Hao Su
摘要
The success of many RL techniques heavily relies on human-engineered dense rewards, which typically demand substantial domain expertise and extensive trial and error. In our work, we propose DrS (Dense reward learning from Stages), a novel approach for learning reusable dense rewards for multi-stage tasks in a data-driven manner. By leveraging the stage structures of the task, DrS learns a high-quality dense reward from sparse rewards and demonstrations if given. The learned rewards can be reused in unseen tasks, thus reducing the human effort for reward engineering. Extensive experiments on three physical robot manipulation task families with 1000+ task variants demonstrate that our learned rewards can be reused in unseen tasks, resulting in improved performance and sample efficiency of RL algorithms. The learned rewards even achieve comparable performance to human-engineered rewards on some tasks. See our project page (https://sites.google.com/view/iclr24drs) for more details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsSenyu Fei, Siyin Wang, Li Ji, Ao Li 等CVPR 2026 · 被引用 28 次
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon TasksMingxuan Yan, Yuping Wang, Zechun Liu, Jiachen LiNeurIPS 2025 · 被引用 4 次
- Cooperative Bargaining Games Without Utilities: Mediated Solutions from Direction OraclesKushagra Gupta, Surya Murthy, Mustafa O. Karabag, Ufuk Topcu 等NeurIPS 2025 · 被引用 2 次
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL ProblemsEgor Cherepanov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 被引用 1 次
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee 等ICLR 2025
它引用的顶会 Paper5
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent 等NeurIPS 2021 · 被引用 106 次
- State Alignment-based Imitation LearningFangchen Liu, Zhan Ling, Tongzhou Mu, Hao SuICLR 2020 · 被引用 103 次
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling 等ICLR 2023 · 被引用 21 次
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 被引用 10 次
- Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard 等CVPR 2020
相关 Paper
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu 等ICML 2025
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 被引用 72 次
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal DistanceYuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman 等ICML 2026 · 被引用 7 次
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen 等NeurIPS 2025 · 被引用 3 次
