DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto
摘要
The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remains challenging to learn and are typically developed for task-specific solutions with online policy learning. To unlock world models' true potential, we argue that they should 1) be trainable on offline, pre-collected trajectories, 2) support test-time behavior optimization, and 3) facilitate task-agnostic reasoning. To this end, we present DINO World Model (DINO-WM), a new method to model visual dynamics without reconstructing the visual world. DINO-WM leverages spatial patch features pre-trained with DI-NOv2, enabling it to learn from offline behavioral trajectories by predicting future patch features. This allows DINO-WM to achieve observational goals through action sequence optimization, facilitating task-agnostic planning by treating goal features as prediction targets. We demonstrate that DINO-WM achieves zero-shot behavioral solutions at test time on six environments without expert demonstrations, reward modeling, or prelearned inverse models, outperforming prior stateof-the-art work across diverse task families such as arbitrarily configured mazes, push manipulation with varied object shapes, and multi-particle scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Cambrian-S: Towards Spatial Supersensing in VideoShusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown 等ICLR 2026 · 被引用 139 次
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero 等NeurIPS 2025 · 被引用 109 次
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung 等ICLR 2026 · 被引用 60 次
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han 等ICLR 2026 · 被引用 38 次
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
相关 Paper
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 被引用 44 次
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider 等ICML 2023 · 被引用 7 次
- AETHER: Geometric-Aware Unified World ModelingHaoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang 等ICCV 2025 · 被引用 9 次
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 被引用 71 次
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone 等NeurIPS 2025 · 被引用 15 次
