ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
Nan Tang, Jing-Cheng Pang, Guanlin Li, Chao Qian, Yang Yu
Abstract
Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on the distance to a target position. However, such precise positional information is often unavailable in real-world visual settings due to sensory and perceptual limitations. In this study, we propose a method that implicitly infers spatial distances through keypoints extracted from images. Building on this, we introduce Reward Learning with Anticipation Model (ReLAM), a novel framework that automatically generates dense, structured rewards from action-free video demonstrations. ReLAM first learns an anticipation model that serves as a planner and proposes intermediate keypoint-based subgoals on the optimal path to the final goal, creating a structured learning curriculum directly aligned with the task's geometric objectives. Based on the anticipated subgoals, a continuous reward signal is provided to train a low-level, goal-conditioned policy under the hierarchical reinforcement learning (HRL) framework with provable sub-optimality bound. Extensive experiments on complex, long-horizon manipulation tasks show that ReLAM significantly accelerates learning and achieves superior performance compared to state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83650a94-b62a-4559-8071-8a9370822bd6Builds on12
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch et al.NeurIPS 2023 · 182 citations
- Video Prediction Models as Rewards for Reinforcement LearningAlejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain et al.NeurIPS 2023 · 117 citations
- Visual Adversarial Imitation Learning using Variational ModelsRafael Rafailov, Tianhe Yu, Aravind Rajeswaran, Chelsea FinnNeurIPS 2021 · 61 citations
Related papers
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji et al.AAAI 2026 · 21 citations
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee et al.ICLR 2025
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 69 citations
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen et al.NeurIPS 2025 · 3 citations
- Training-Free Generation of Temporally Consistent Rewards from VLMsYinuo Zhao, Jiale Yuan, Zhiyuan Xu, Xiaoshuai Hao et al.ICCV 2025 · 1 citation
