TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
Yuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman, Yang Gao
Abstract
Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies the degree to which actions advance the system toward task completion over time. We present TimeRewarder, a simple yet effective reward learning method that derives progress estimation signals from passive videos, including robot demonstrations and human videos, by modeling temporal distances between frame pairs. We then demonstrate how TimeRewarder can supply step-wise proxy rewards to guide reinforcement learning. In our comprehensive experiments on ten challenging Meta-World tasks, we show that TimeRewarder dramatically improves RL for sparse-reward tasks, achieving nearly perfect success in 9/10 tasks with only 200,000 environment interactions per task. This approach outperforms previous methods and even the manually designed environment dense reward on both the final success rate and sample efficiency. Moreover, we show that TimeRewarder can exploit real-world human videos, highlighting its potential as a scalable approach to rich reward signals from diverse video sources. Project page: timerewarder.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9509cd8-0d1f-4005-86ca-9c94ee507036Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- Video Prediction Models as Rewards for Reinforcement LearningAlejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain et al.NeurIPS 2023 · 117 citations
Related papers
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang et al.ICCV 2025 · 13 citations
- GUIDE: Real-Time Human-Shaped AgentsLingyu Zhang, Zhengran Ji, Nicholas R. Waytowich, Boyuan ChenNeurIPS 2024 · 9 citations
- Generalizable Imitation Learning from Observation via Inferring Goal ProximityYoungwoon Lee, Andrew Szot, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 64 citations
- Robot Policy Learning with Temporal Optimal Transport RewardYuwei Fu, Haichao Zhang, Di Wu, Wei Xu et al.NeurIPS 2024 · 13 citations
- Imitation Learning from a Single Temporally Misaligned VideoWilliam Huey, Huaxiaoyue Wang, Anne Wu, Yoav Artzi et al.ICML 2025
