Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online Refinement
Tewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang, Michael Maire, Matthew R. Walter
Abstract
We present Progressor, a novel framework that learns a task-agnostic reward function from videos, enabling policy training through goal-conditioned reinforcement learning (RL) without manual supervision. Underlying this reward is an estimate of the distribution over task progress as a function of the current, initial, and goal observations that is learned in a self-supervised fashion. Crucially, Progressor refines rewards adversarially during online training by pushing back predictions for out-of-distribution observations in order to mitigate distribution shift inherent in non-expert observations. Utilizing this progress prediction as a dense reward together with an adversarial pushback, we show that Progressor enables robots to learn complex behaviors without any external supervision. Pretrained on large-scale egocentric human video from EPICKITCHENS, Progressor requires no fine-tuning on indomain task-specific data for generalization to real-robot offline RL under noisy demonstrations, outperforming contemporary methods that provide dense visual reward for robotic learning. Our findings highlight the potential of Progressor for scalable robotic applications where direct action labels and task-specific rewards are not readily available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49fe0de7-a913-42a0-bc51-fa1e215b8dbeCited by top-tier papers2
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal DistanceYuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman et al.ICML 2026 · 7 citations
- MVR: Multi-view Video Reward Shaping for Reinforcement LearningLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang et al.ICLR 2026
Builds on9
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
Related papers
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
- Generalizable Imitation Learning from Observation via Inferring Goal ProximityYoungwoon Lee, Andrew Szot, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 64 citations
- Shaping embodied agent behavior with activity-context priors from egocentric videoTushar Nagarajan, Kristen GraumanNeurIPS 2021 · 23 citations
- Video Prediction Models as Rewards for Reinforcement LearningAlejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain et al.NeurIPS 2023 · 117 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
