Grounding Video Models to Actions through Goal Conditioned Exploration
Yunhao Luo, Yilun Du
Abstract
Large video models, pretrained on massive amounts of Internet video, provide a rich source of physical knowledge about the dynamics and motions of objects and tasks. However, video models are not grounded in the embodiment of an agent, and do not describe how to actuate the world to reach the visual states depicted in a video. To tackle this problem, current methods use a separate vision-based inverse dynamic model trained on embodiment-specific data to map image states to actions. Gathering data to train such a model is often expensive and challenging, and this model is limited to visual settings similar to the ones in which data are available. In this paper, we investigate how to directly ground video models to continuous actions through self-exploration in the embodied environment -using generated video states as visual goals for exploration. We propose a framework that uses trajectory level action generation in combination with video guidance to enable an agent to solve complex tasks without any external supervision, e.g., rewards, action labels, or segmentation masks. We validate the proposed approach on 8 tasks in Libero, 6 tasks in MetaWorld, 4 tasks in Calvin, and 12 tasks in iThor Visual Navigation. We show how our approach is on par with or even surpasses multiple behavior cloning baselines trained on expert demonstrations while without requiring any action annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d9ef0c4-44de-4708-9c29-60c705fdfedfCited by top-tier papers8
- ORV: 4D Occupancy-centric Robot Video GenerationXiuyu Yang, Bohan Li, Shaocong Xu, Nan Wang et al.CVPR 2026 · 19 citations
- Medical World ModelYijun Yang, Zhao-Yang Wang, Qiuping Liu, Shuwen Sun et al.ICCV 2025 · 7 citations
- Goal-Driven Reward by Video Diffusion Models for Reinforcement LearningQi Wang, Mian Wu, Yuyang Zhang, Mingqi Yuan et al.CVPR 2026 · 2 citations
- Translating Flow to Policy via Hindsight Online ImitationYitian Zheng, Zhangchen Ye, Weijun Dong, Shengjie Wang et al.ICLR 2026 · 2 citations
- Empowering World Models with Reflection for Embodied Video PredictionXiaowei Chi, Chun-Kai Fan, Hengyuan Zhang, Xingqun Qi et al.ICML 2025
Builds on27
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
Related papers
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng et al.ICLR 2026 · 18 citations
- VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot ManipulationYoupeng Wen, Junfan Lin, Yi Zhu, Jianhua Han et al.NeurIPS 2024 · 63 citations
- Self-Improving Loops for Visual Robotic PlanningCalvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du et al.ICLR 2026 · 4 citations
