Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
Shivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain, Svetlana Lazebnik, Yunzhu Li
摘要
This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as pouring, wiping, and mixing--purely by imitating AI-generated videos, without requiring any physical demonstrations or robot-specific training. Given a language command and an initial scene image, a video diffusion model generates potential demonstration videos, and a vision-language model (VLM) automatically filters out results that do not follow the command. A 6D pose tracker then extracts object trajectories from the video, and the trajectories are retargeted to the robot in an embodiment-agnostic fashion. Through extensive real-world evaluations, we show that filtered generated videos are as effective as real demonstrations, and that performance improves with generation quality. We also show that relying on generated videos outperforms more compact alternatives such as keypoint prediction using VLMs, and that strong 6D pose tracking outperforms other ways to extract trajectories, such as dense feature point tracking. These findings suggest that videos produced by a state-of-the-art off-the-shelf model can offer an effective source of supervision for robotic manipulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu 等CVPR 2026 · 被引用 87 次
- ViPRA: Video Prediction for Robot ActionsSandeep Kumar Routray, Hengkai Pan, Unnat Jain, Shikhar Bahl 等ICLR 2026 · 被引用 30 次
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesGabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht 等ICLR 2026 · 被引用 10 次
- Unifying Stacking and Cascading for Efficient Ensemble InferenceAshwin Colaço, Sharad Mehrotra, Michael De Lucia, Kevin Hamlen 等ICML 2026
- Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot LearningYinan Deng, Kejia Hu, Ye Chen, Jianyu Dou 等CVPR 2026
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
相关 Paper
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- 6D Object Pose Tracking in Internet Videos for Robotic ManipulationGeorgy Ponimatkin, Martin Cífka, Tomás Soucek, Médéric Fourmy 等ICLR 2025
- NIL: No-data Imitation LearningMert Albaba, Chenhao Li, Markos Diomataris, Omid Taheri 等CVPR 2026
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang 等ICLR 2026
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys 等CVPR 2025
