Learning Universal Policies via Text-Guided Video Generation
Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, Pieter Abbeel
摘要
A goal of artificial intelligence is to construct an agent that can solve a wide variety of tasks. Recent progress in text-guided image synthesis has yielded models with an impressive ability to generate complex novel images, exhibiting combinatorial generalization across domains. Motivated by this success, we investigate whether such tools can be used to construct more general-purpose agents. Specifically, we cast the sequential decision making problem as a text-conditioned video generation problem, where, given a text-encoded specification of a desired goal, a planner synthesizes a set of future frames depicting its planned actions in the future, after which control actions are extracted from the generated video. By leveraging text as the underlying goal specification, we are able to naturally and combinatorially generalize to novel goals. The proposed policy-as-video formulation can further represent environments with different state and action spaces in a unified space of images, which, for example, enables learning and generalization across a variety of robot manipulation tasks. Finally, by leveraging pretrained language embeddings and widely available videos from the internet, the approach enables knowledge transfer through predicting highly realistic video plans for real robots 2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper165
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson 等ICLR 2024 · 被引用 399 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- Zero-Shot Robotic Manipulation with Pre-Trained Image-Editing Diffusion ModelsKevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke 等ICLR 2024 · 被引用 284 次
- UniControl: A Unified Diffusion Model for Controllable Visual Generation In the WildCan Qin, Shu Zhang, Ning Yu, Yihao Feng 等NeurIPS 2023 · 被引用 250 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun 等ICLR 2024 · 被引用 181 次
- Solving New Tasks by Adapting Internet Video KnowledgeCalvin Luo, Zilai Zeng, Yilun Du, Chen SunICLR 2025
- RoboDreamer: Learning Compositional World Models for Robot ImaginationSiyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li 等ICML 2024 · 被引用 140 次
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia 等ICLR 2024 · 被引用 161 次
