Solving New Tasks by Adapting Internet Video Knowledge
Calvin Luo, Zilai Zeng, Yilun Du, Chen Sun
Abstract
Video generative models demonstrate great promise in robotics by serving as visual planners or as policy supervisors. When pretrained on internet-scale data, such video models intimately understand alignment with natural language, and can thus facilitate generalization to novel downstream behavior through textconditioning. However, they may not be sensitive to the specificities of the particular environment the agent inhabits. On the other hand, training video models on in-domain examples of robotic behavior naturally encodes environmentspecific intricacies, but the scale of available demonstrations may not be sufficient to support generalization to unseen tasks via natural language specification. In this work, we investigate different adaptation techniques that integrate in-domain information with large-scale pretrained video models, and explore the extent to which they enable novel text-conditioned generalization for robotic tasks, while also considering their independent data and resource considerations. We successfully demonstrate across robotic environments that adapting powerful video models with small scales of example data can successfully facilitate generalization to novel behaviors. In particular, we present a novel adaptation strategy, termed Inverse Probabilistic Adaptation, that not only consistently achieves strong generalization performance across robotic tasks and settings, but also exhibits robustness to the quality of adaptation data, successfully solving novel tasks even when only suboptimal in-domain demonstrations are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da4388d6-0474-473a-87cb-e8373e25a2d1Cited by top-tier papers7
- ViPRA: Video Prediction for Robot ActionsSandeep Kumar Routray, Hengkai Pan, Unnat Jain, Shikhar Bahl et al.ICLR 2026 · 30 citations
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denosing Diffusion ProcessJiayi Chen, Wenxuan Song, Pengxiang Ding, Ziyang Zhou et al.ICLR 2026 · 20 citations
- Learning Interactive World Model for Object-Centric Reinforcement LearningFan Feng, Phillip Lippe, Sara MagliacaneNeurIPS 2025 · 13 citations
- Goal Force: Teaching Video Models To Accomplish Physics-Conditioned GoalsNate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo et al.CVPR 2026 · 12 citations
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli et al.NeurIPS 2025 · 8 citations
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
- Probabilistic Adaptation of Black-Box Text-to-Video ModelsSherry Yang, Yilun Du, Bo Dai, Dale Schuurmans et al.ICLR 2024 · 8 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- Self-Improving Loops for Visual Robotic PlanningCalvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du et al.ICLR 2026 · 4 citations
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li et al.ICCV 2025 · 5 citations
