VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
Zhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao, Bingyi Kang, Jiashi Feng, Xiaojie Jin
摘要
This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language models (LLMs). We develop VideoWorld, an auto-regressive video generation model trained on unlabeled video data, and test its knowledge acquisition abilities in video-based Go and robotic control tasks. Our experiments reveal two key findings: (1) video-only training provides sufficient information for learning knowledge, including rules, reasoning and planning capabilities, and (2) the representation of visual change is crucial for knowledge acquisition. To improve both the efficiency and efficacy of this process, we introduce the Latent Dynamics Model (LDM) as a key component of VideoWorld. Remarkably, VideoWorld reaches a 5-dan professional level in the Video-GoBench with just a 300-million-parameter model, without relying on search algorithms or reward mechanisms typical in reinforcement learning. In robotic tasks, VideoWorld effectively learns diverse control operations and generalizes across environments, approaching the performance of oracle models in CALVIN and RLBench. This study opens new avenues for knowledge acquisition from visual data, with all code, data, and models open-sourced for further research. * Following prior works in AI and knowledge representation [5, 51] which view 'knowledge' as extending beyond factual information, we use 'knowledge' to broadly refer to a model's learned rules, reasoning, and planning abilities necessary for task completion. For better readability and clarity, these terms are used interchangeably in specific contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningSiyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang 等ACL 2025 · 被引用 16 次
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh 等ICLR 2026 · 被引用 12 次
- Video-GPT via Next Clip DiffusionShaobin Zhuang, Zhipeng Huang, Ying Zhang, Fangyikang Wang 等ICLR 2026 · 被引用 9 次
- Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal InconsistencyJiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang 等ACL 2025 · 被引用 2 次
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationXinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- VideoWorld 2: Learning Transferable Knowledge from Real-world VideosZhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo 等CVPR 2026 · 被引用 9 次
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang 等ICML 2025
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li 等ICCV 2025 · 被引用 5 次
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik 等ICML 2026 · 被引用 96 次
- Grounding Video Models to Actions through Goal Conditioned ExplorationYunhao Luo, Yilun DuICLR 2025
