Learning World Models for Interactive Video Generation
Taiye Chen, Xun Hu, Zihan Ding, Chi Jin
Abstract
Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities due to two main challenges: compounding errors and insufficient memory mechanisms. We enhance image-to-video models with interactive capabilities through additional action conditioning and autoregressive framework, and reveal that compounding error is inherently irreducible in autoregressive video generation, while insufficient memory mechanism leads to incoherence of world models. We propose video retrieval augmented generation (VRAG) with explicit global state conditioning, which significantly reduces long-term compounding errors and increases spatiotemporal consistency of world models. In contrast, naive autoregressive generation with extended context windows and retrieval-augmented generation prove less effective for video generation, primarily due to the limited in-context learning capabilities of current video models. Our work illuminates the fundamental challenges in video world models and establishes a comprehensive benchmark for improving video generation models with internal world modeling capabilities. Project page: https://sites.google.com/view/vrag .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu et al.ICML 2026 · 108 citations
- Plenoptic Video GenerationXiao Fu, Shitao Tang, Min Shi, Xian Liu et al.CVPR 2026 · 10 citations
Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li et al.AAAI 2026 · 1 citation
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video UnderstandingXiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed ElhoseinyNeurIPS 2025 · 37 citations
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu et al.NeurIPS 2025 · 145 citations
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video BenchmarkSeng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao et al.CVPR 2026
- LIVE: Long-horizon Interactive Video World ModelingJunchao Huang, Ziyang Ye, Xinting Hu, Tianyu He et al.ICML 2026 · 12 citations
