Grounding Video Models to Actions through Goal Conditioned Exploration
Yunhao Luo, Yilun Du
摘要
Large video models, pretrained on massive amounts of Internet video, provide a rich source of physical knowledge about the dynamics and motions of objects and tasks. However, video models are not grounded in the embodiment of an agent, and do not describe how to actuate the world to reach the visual states depicted in a video. To tackle this problem, current methods use a separate vision-based inverse dynamic model trained on embodiment-specific data to map image states to actions. Gathering data to train such a model is often expensive and challenging, and this model is limited to visual settings similar to the ones in which data are available. In this paper, we investigate how to directly ground video models to continuous actions through self-exploration in the embodied environment -using generated video states as visual goals for exploration. We propose a framework that uses trajectory level action generation in combination with video guidance to enable an agent to solve complex tasks without any external supervision, e.g., rewards, action labels, or segmentation masks. We validate the proposed approach on 8 tasks in Libero, 6 tasks in MetaWorld, 4 tasks in Calvin, and 12 tasks in iThor Visual Navigation. We show how our approach is on par with or even surpasses multiple behavior cloning baselines trained on expert demonstrations while without requiring any action annotations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ORV: 4D Occupancy-centric Robot Video GenerationXiuyu Yang, Bohan Li, Shaocong Xu, Nan Wang 等CVPR 2026 · 被引用 19 次
- Medical World ModelYijun Yang, Zhao-Yang Wang, Qiuping Liu, Shuwen Sun 等ICCV 2025 · 被引用 7 次
- Goal-Driven Reward by Video Diffusion Models for Reinforcement LearningQi Wang, Mian Wu, Yuyang Zhang, Mingqi Yuan 等CVPR 2026 · 被引用 2 次
- Translating Flow to Policy via Hindsight Online ImitationYitian Zheng, Zhangchen Ye, Weijun Dong, Shengjie Wang 等ICLR 2026 · 被引用 2 次
- Empowering World Models with Reflection for Embodied Video PredictionXiaowei Chi, Chun-Kai Fan, Hengyuan Zhang, Xingqun Qi 等ICML 2025
它引用的顶会 Paper27
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
相关 Paper
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 被引用 98 次
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma 等ICML 2026 · 被引用 2 次
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng 等ICLR 2026 · 被引用 18 次
- VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot ManipulationYoupeng Wen, Junfan Lin, Yi Zhu, Jianhua Han 等NeurIPS 2024 · 被引用 63 次
- Self-Improving Loops for Visual Robotic PlanningCalvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du 等ICLR 2026 · 被引用 4 次
